Hacker Newsnew | past | comments | ask | show | jobs | submit | ekojs's commentslogin

Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.

[0]: https://x.com/EpochAIResearch/status/2095602754282783108


> In our evaluations, Kimi K3 delivers frontier-level performance. Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol. For the complete benchmark results, see our tech blog. The full model weights of Kimi K3 will be released in the coming days. More details on the architecture, training, and evaluation will be published together with the Kimi K3 technical report.

> K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600.

> On AA-Briefcase, Kimi K3 scores 1527, ranking second among all models — behind only Claude Fable 5 Max and ahead of GPT-5.6 Sol Max (1495). AA-Briefcase is a private agentic knowledge-work benchmark developed by Artificial Analysis to evaluate frontier agentic capability in long-horizon knowledge work.

Really good benchmark score it seems. Maybe another DeepSeek moment right here.


> its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol

Pretty sure ranking “second” to two others means ranking third.


Charitably, you could read this as "its overall intelligence [is in a class that] ranks second only to [that of]..."


This is actually what's meant but this bikeshed has been built for yak shaving.


Since we’re bikeshedding: that’s not what yak shaving means.


Since we’re hair splitting: that’s not what bike shedding means. :-D


What does it look like I'm doing??


yakshedding? bikeshaving? IDK


Yeah, bad wording it seems. Though a charitable interpretation is that Fable 5 and GPT 5.6 Sol are joint 1st place in the measurement.


Doesn’t matter, the next one is still third.


DENSE_RANK() vs RANK() claims another victim


girl, we get it, you can count. re-stating your point does not a conversation make.


If there are two folks standing at gold, nobody gets the silver medal.


But linearizing an equal magnitude quantities by alphabet priority would be unfair. Magnitude is the important quantity here.


"Ranks second" is their statement. What is it's rank, in your opinion?


frontier vs "not quite" :D


While you are technically correct, in English it’s perfectly fine to say it this way as well.

“Second only” here has meaning “next after”, not “number two”.


Yes. "Second to" takes a set as an argument in English. Even the empty set works!

England is second to none.


So... France took second to England and Argentina?


France’s football team is second only to England’s and Argentina’s.

It’s a miracle that in language same words have different meanings depending on context. If this wouldn’t be the case we could have hardcoded NLP algorithmically without inventing these expensive LLMs!


Second group essentially is how you have to think of it


That’s not what second means in this context in English, and it’s incorrect to use it that way. This is because for something to be second there must have been something in first and only first, and so on; in this case there was a first and a second already, and you cannot amalgamate then because they didn’t tie (and even if they did, they’d be 1 and 2). Both logically and grammatically, it’s incorrect.


You're both logically and grammatically wrong. You could even ask an LLM to explain the meaning of the phrase to you if you don't believe that.


Either think and write for yourself or stay silent next time. It'd be infinitely better than telling another person to use an LLM to understand something you yourself don't understand and are too lazy to try to figure out.

I wish you the best.


Hah, I had expected this knee-jerk response, but kinda hoped you'd avoid this pitfall. Alas.

See, I could tell you that in English, "second to" is a construct that usually means "next to" or "inferior to" and has nothing to do with "being in second place", and that if it did, it would make the popular construct "second only to" completely redundant. But others already did that in sibling comments before me, and you could just respond with "you're wrong" anyway, so what's the point? Pointing to an LLM is, of course, often a lazy and unhelpful cop out from the discussion, but in this particular case it's pointing you to a dataset that's explicitly about extracting meaning and finding relationships between phrases in languages - so you don't have to trust me or anyone else that this phrase is actually being used in this particular way, you can find it out yourself based on enormous training datasets illegally collected from all over the Internet.


If what you are saying were true, I could rank anything high by simply putting everything actually higher than it in rank into one named set and then turn around and say that the thing I want to highly rank comes in 2nd only to that first set of n items.


You haven't read a word of what I wrote, have you?


Not if the others tie for first place.


Still third even then.


I think what's implied here, in colloquial terms, is that it's in the second tier.


Except saying second tier would be bad marketing, so they decided to change the meaning of words instead.


okay, so let’s get this straight: even though there are seemingly quite a few people here that clearly understand what is being said. However, the fact that YOU specifically either genuinely do not understand this / have never come across this before, or are being intentionally difficult because of some philosophical disagreement, feel that you can unilaterally assert that they’re “redefining words”?

I don’t know if this is genuinely your first day on Earth or something, but if you’re trying to parse English like a programming language then you’re not only making things hard on yourself, but also 99% of people you’ll ever speak to.


Which is still great because it means neither of the two best financed labs in the world manage to produce even two models themselves that would beat Kimi K3.


> > K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600.

This is the same benchmark where Sonnet 5 outperforms Opus 4.8 max.

Like all model releases, the benchmarks aren't going to tell the whole story. All of the open weight models come with amazing benchmark results now. It's hard to believe anything other than that the benchmarks are leaking into (or intentionally included) into training data.


Possible, but pay-as-you-go Hy3 / DeepSeek v4 Pro / MiMo v2.5 Pro (from respective vendors) are genuinely good enough as daily drivers, given the costs (especially, low prices for input cache, which usually makes up 70%+ of total input for agentic workflows). I put in $10 in DeepSeek & Xiaomi MiMo, and I've barely used $1 each, in a week of coding work.

Coding Plans by MiniMax ($20/mo for 1.7b tokens) and Z.ai (~$30/week use for $17/mo) are also tremendous value for money.


Sonnet 5 does beat Opus 4.8 on several benchmarks. It just costs more and takes longer.

(On several other benchmarks, it costs more, takes longer, and does worse.)


i’ll never really understand this comment. why would labs do this if they know private benchmark evals will come out in the next week?


> Maybe another DeepSeek moment right here.

Surely not... What made DeepSeek disruptive was that the cost was 10X lower.

In this case, the cost is about 2X lower the Sol I think?

At 2X, you're pretty close to the error margins due to token efficiency etc...

I'd say this is "on trend" for open models catching up to frontier labs, but its not a "change in the trend" like DeepSeek was IMO.


It was also disruptive because it was open weight, meaning anyone and their dog could theoretically compete with the frontier labs for their inference revenue.

The frontier labs need to recoup a huge amount of cash to cover their model development costs, and justify their valuations. That’s plausible when they’re only ones capable of selling inference on these models, it a lot less plausible when models themselves become cheap commodities, and you’re just competing on your ability to provide compute. Anthropic and OpenAI can’t compete with people like AWS on that front.


It's different, but similar. If they release the weights, then we have a Fable / frontier model people can tinker with. Either way, it's still quite impressive and knocked a US company out of the top three (google). How long before China dominates the top-10 (if they don't already) or the #1 model?


Moonshot announced they will be making weights available by July 27

https://twitter.com/Kimi_Moonshot/status/2077830229968683203


cost has nothing to do with why deepseek was disruptive, the fact that it means there is zero moat around anthropic or openai is what's disruptive about it. it means in the mid-term LLMs will be commoditized and customers will flock to the cheapest inference wherever they can find it. there's no reason to stick to the "frontier" labs


> cost has nothing to do with it

> customers will flock to the cheapest inference


if deepseek cost twice as much to train it would prove the same thing: the american companies have no monopoly on state of the art llms, and commoditization is happening


If AA is to be believed then per-task it is about the same cost as Sol. Agree that it's very different from DeepSeek v4 Pro, which is ~15x cheaper than K3.

https://artificialanalysis.ai/models/kimi-k3#price-cost


DeepSeek didn’t really change any trends though, unless you count the stock market.

It was impressive work, but models were commoditizing and inference costs were dropping rapidly already. They were neither the first nor the last 10x optimization, from what I’ve seen.


If you know of any other 10x optimisations currently, please let me know! I'm in the market for a model that's a tenth the price of a frontier model at the same level of quality.


You do understand that the "frontier" people are usually talking about is the cost-intelligence frontier right?

By definition there is no model that is both cheaper and as intelligent or better than another on the frontier.


OK, let me be more precise: If you know of a frontier model that's ten times cheaper than the previous frontier model at that level of intelligence, please let me know, I'm in the market for one.

Is that better?


Its just a single benchmark, but Luna 5.6 xhigh scores within the margin of error the same as Opus 4.8 max on DeepSWE for 8x cheaper. Luna max is quite a bit higher than Opus and still 4x cheaper


Oh really? That's great to know, thanks! I didn't realise Luna was that good.


To be fair the stock market is a big one


In my experience, the Chinese models are much more benchmaxxed than their frontier lab competitors, so I'm taking these results with a fairly large helping of salt.


That’s an interesting way to say you’re third. I’m only second to the ten other runners on my local Strava segments.


> In our evaluations, Kimi K3 delivers frontier-level performance

What page does that come from? I'm having trouble tracking it down.


It was on the page linked in the top comment, but it's been removed.


Where are you seeing this write up?


I copied that from https://platform.kimi.ai/docs/guide/kimi-k3-quickstart but it seems they updated the page to remove the benchmark score now.


What hardware are they using to run the benchmarks?


Where is this from?


Seems like the only good thing about 3.5 Flash is its speed. Not cost-competitive or benchmark-leading by any means.


Not disagreeing with your argument, but:

> If you want a good dense model, use qwen3.6 27B instead, speed will be up, and if you don't take my word for it being smarter, take openrouter's prices of it against the bigger, slower and less memory-efficient gemma do the talking.

Don't know if this is the correct read. I think those providers are simply taking cue from Alibaba's first-party pricing for the 27B Dense. It's kinda overpriced imo. Perhaps it can be explained by how 'reasoning-inefficient' (relative to frontier models or even Gemma) the Qwen models are and longer sequence lengths are expensive to serve.


I feel like if I had the infrastructure and saw that there is a huge interest in the model, i'd just undercut alibaba's prices a little harder to grab all the consumers. I am sure that the providers have done the math and found that there is a reason not to do this (compute-bound if too many users?), but the delta is very stark, especially for output. Last I checked the cheapest 27b on openrouter was 2$ out vs 0.38$ for the 31b.

But I do agree that the openrouter prices aren't a strong signal and probably should have worded it a little better. It's just a really stark and 'in your eyes' gap.


> HTTP is just not a good transport for streaming LLM tokens and for building async agentic applications

I don't know if I agree if this is a problem with SSE or HTTP. Something like a Redis Streams-backed SSE would solve most of the 'challenges' presented in the post.


As this is a dense model and it's pretty sizable, 4-bit quantization can be nearly lossless. With that, you can run this on a 3090/4090/5090. You can probably even go FP8 with 5090 (though there will be tradeoffs). Probably ~70 tok/s on a 5090 and roughly half that on a 4090/3090. With speculative decoding, you can get even faster (2-3x I'd say). Pretty amazing what you can get locally.


> As this is a dense model and it's pretty sizable, 4-bit quantization can be nearly lossless

The 4-bit quants are far from lossless. The effects show up more on longer context problems.

> You can probably even go FP8 with 5090 (though there will be tradeoffs)

You cannot run these models at 8-bit on a 32GB card because you need space for context. Typically it would be Q5 on a 32GB card to fit context lengths needed for anything other than short answers.


I just loaded up Qwen3.6 27B at Q8_0 quantization in llama.cpp, with 131072 context and Q8 kv cache:

  build/bin/llama-server \
    -m ~/models/llm/qwen3.6-27b/qwen3.6-27B-q8_0.gguf \
    --no-mmap \
    --n-gpu-layers all \
    --ctx-size 131072 \
    --flash-attn on \
    --cache-type-k q8_0 \
    --cache-type-v q8_0 \
    --jinja \
    --no-mmproj \
    --parallel 1 \
    --cache-ram 4096 -ctxcp 2 \
    --reasoning on \
    --chat-template-kwargs '{"preserve_thinking": true}'
Should fit nicely in a single 5090:

  self    model   context   compute
  30968 = 25972 +    4501 +     495
Even bumping up to 16-bit K cache should fit comfortably by dropping down to 64K context, which is still a pretty decent amount. I would try both. I'm not sure how tolerant Qwen3.5 series is of dropping K cache to 8 bits.


> You cannot run these models at 8-bit on a 32GB card because you need space for context

You probably can actually. Not saying that it would be ideal but it can fit entirely in VRAM (if you make sure to quantize the attention layers). KV cache quantization and not loading the vision tower would help quite a bit. Not ideal for long context, but it should be very much possible.

I addressed the lossless claim in another reply but I guess it really depends on what the model is used for. For my usecases, it's nearly lossless I'd say.


Turboquant on 4bit helps a lot as well for keeping context in vram, but int4 is definitely not lossless. But it all depends for some people this is sufficient


4-bit quantization is almost never lossless especially for agentic work, it's the lowest end of what's reasonable. It's advocated as preferable to a model with fewer parameters that's been quantized with more precision.


Yeah, figure the 'nearly lossless' claim is the most controversial thing. But in my defense, ~97% recovery in benchmarks is what I consider 'nearly lossless'. When quantized with calibration data for a specialized domain, the difference in my internal benchmark is pretty much indistinguishable. But for agentic work, 4-bit quants can indeed fall a bit short in long-context usecase, especially if you quantize the attention layers.


4-bit quantization is not applied to all layers, some are kept 8/16-bit.


That seems awfully speculative without at least some anecdata to back it up.


Sure, go get some.

This isn't the first open-weight LLM to be released. People tend to get a feel for this stuff over time.

Let me give you some more baseless speculation: Based on the quality of the 3.5 27B and the 3.6 35B models, this model is going to absolutely crush it.


Not at all, I actually run ~30B dense models for production and have tested out 5090/3090 for that. There are gotchas of course, but the speed/quality claims should be roughly there.


Seems pretty widespread. We got mistakenly charged for ~$800 over the weekend.

Other Sources:

[0]: https://aistudio.google.com/status

[1]: https://www.reddit.com/r/GeminiAI/comments/1mycmtk/google_cl...

[2]: https://www.reddit.com/r/GeminiAI/comments/1myg04q/gemini_25...


> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved

> We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results officially graded and certified by IMO coordinators and experts, receiving the first official gold-level performance grading for an AI system!

From https://x.com/demishassabis/status/1947337620226240803

Was OpenAI simply not coordinating with the IMO Board then?


Yes, there have been multiple (very big) hints dropped by various people that they had no official cooperation.


I think this is them not being confident enough before the event, so they don't wanna be shown a worse result than competitors. By being private they can obviously not publish anything if it didn't work out.


They shot themselves in the foot by not showing the confidence that Google did.


As not-so-subtly hinted at by Terry Tao.

Its a great way to do PR but its a garbage way to to science.


True, but openai definitely isn't trying to do public research on science, they are all about money now.


Thats not a contentious statement. Its still a pathetic way to behave at a kids competition no less.


This reminds me of when OpenAI made a splash (ages ago now) by beating the world's best Dota 2 teams using a RL model.

...Except they had to substantially bend the rules of the game (limiting the hero pool, completely changing/omitting certain mechanics) to pull this off. So they ended up beating some human Dota pros at a psuedo-Dota custom game, which was still impressive, but a very much watered-down result beneath the marketing hype.

It does seem like Money+Attention outweigh Science+Transparency at OpenAI, and this has always been the case.


Limiting the hero pool was fair I'd say. If you can prove RL works on one hero, it's fairly certain it would work on other heroes. All of them at once? Maybe run into problems. But anyway you'd need orders of magnitude more compute so I'd say that was fair game.


It's not even close to the same game as Dota. Limiting the hero (and item) pool so drastically locks off many strategies and counters. It's a bit hard to explain if you haven't played, but full Dota has many more tools and much more creativity than the reduced version on display. The behavior does not evidently "scale up", in the same way that the current SotA of AI art and writing won't evidently replace top-level humans.

I'd never say it's impossible, but the job wasn't finished yet.


That's akin to saying it's okay to remove Knights, or castling, or en passant from chess because they have a complicated movement mechanic that the AI can't handle as well.

Hero drafting and strategy is a major aspect of competitive Dota 2.


> Was OpenAI simply not coordinating with the IMO Board then?

You are still surprised by sama@'s asinineness? You must be new here.


When your goal is to control as much of the world's money as possible, preferably all of it, then everyone is your enemy, including high school students.


How dare those high school students use their brains to compete with ChatGPT and deny the shareholders their value?


I am still surprised many people trust him. The board's (justified) decision to fire him was so awfully executed that it lead to him having even more slack


Maybe not a popular sentiment here on HN but I cancelled my Kagi subscription (9+ months) just recently. Increasingly, most of my queries/search have been through LLMs and Google search is just fine (and even better for restaurants, places, and the like). I don't think the improved search experience is worth the subscription anymore.


In Kagi, you can just ad a "?" to your query and get an instant answer, a la LLMs.


Or !ai to route it to kagi.com/assistant with your default model/agent to respond with kagi search results


https://status.cloud.google.com/incidents/ow5i3PPK96RduMcb1S...

> Multiple GCP products are experiencing impact due to Identity and Access Management Service Issue

IAM issue huh. The post-mortem should be interesting at least.


Ha. With all this soviet style euphemism I rather read the onion instead.


It’s not a euphemism - every outage, including the 99.9% that don’t end up on HN gets a postmortem document written about it, which is almost always a fascinating discussion of the technical, cultural and organisational situation that led to an unexpected bad thing happening.

Even a few years ago senior management knew to stay the fuck out except for asking for more info.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: