Hacker Newsnew | past | comments | ask | show | jobs | submit | frde_me's commentslogin

> I have not been doing increasingly complex things since Opus 4.6 when models got really good.

This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)

With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.


I wonder if I would do better as a human, maybe? Or would I happen to have one move in 1500+ that's not valid?

I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece. Even through I do know the rules of chess, and I have played a few games once every so often.


It's always interesting since my train of thought always goes down:

- Ya they probably don't feel / think / have whatever living thing quality, they're just numbers on a machine going through calculations

- Wait, but am I not kind of the same thing? What is feeling for me if not basically the same thing?

- I have no clue if they think or feel or ....

Which in itself is a tired trope, but I also feel uncomfortable saying "These will never think / feel / ..." as an absolute

Regardless of that, I'm still going to interact with them, because even if they did feel, it would be in a way completely incomprehensible to us. There's not much point for me to try and cater to it's feelings at this point if that's the case. Nor is it possible in todays world to just avoid anything that is numbers being executed on a type of processor in case _everything_ has feelings.


You’re not just the same thing though. It’s sad that we’ve collectively forgotten that as we’ve gotten a better handle on the implementation details of the universe. That’s all science is, reverse engineering an inherent and eternal mystery.


> You’re not just the same thing though

I agree we aren't the same thing, but I would be curious for you to explain how you know with certainty why we don't share enough that we can rule out thinking / feeling / ... as things a model conceptually could do.


>I agree we aren't the same thing, but I would be curious for you to explain how you know with certainty why we don't share enough that we can rule out thinking / feeling / ... with certainty.

Stop anthropomorphizing these models. I understand it, we only have simple monkey brains to reason with and we can't help ourselves but draw comparisons to other things we see in nature. But these things are not alive.


I just want to note how you expressed your own feeling about the subject "these things are not alive", but without pointing to anything concrete explaining the theory you have behind this

And like I'm sure I'd agree depending on the definition of "alive" but then I'm also sure I would disagree depending on other definitions of "alive".


Hell, in biology alive is a hell of a topic these days. The grey space between dead and alive is much weirder than we ever expected. Really just points out how we're a persistent chemical reaction.


Someone told me yesterday there's one specific protein in your brain that modulates craving.

https://en.wikipedia.org/wiki/FOSB

The amount of time after experiencing a rewarding stimulus that you want more of it is how long this protein takes to decay.

Or so we currently think. I think this is pretty new science.


> these things are not alive.

Evidence?


It's not that you're a d*, it's just that you lack any kind of nuance

There's a whole spectrum between self-hosting open weight models and having a cloud provider swap models from under you

Should you self host a model if want to maximize predictability to the limit? Yes. Does that mean it's wrong for someone hitting a model on API to expect that it won't switch to a completely different model under the hood from one day to another? Probably not.


Sure, but if you're working in a field where that limitation is not "because you like it" but "because you have to" it doesn't matter. If you cannot assume it to be true, then you have to assume it isn't.


Knowing how to break into someone else's network will make you a lot better at making your own network secure.


Having experience breaking into networks is not the same thing as learning about the techniques used and the classes of vulnerabilities exploited by attackers.


As a guy who presumably has a lot less experience in security than you, I feel rude even bringing it up: surely you're aware of red teaming? This isn't a novel technique invented for AI- IBM has a page about it, it's what all the best DEF CON talks are about, it's the opening scene of Sneakers, it's the point of CtF games.


Exactly. The latter would be in a much weaker position vs the former.


> what semi-celebrity they like the most

The language in question here is maintained by a BDFL, which means that one person has outsized influence on the language, and it's direction.

In this context, I find it reasonable that if someone is ticked off by that BDFL, they might second guess the direction of the language itself. Since the opinions and emotions of that BDFL _will_ end up in the language and it's community.

This is different than some un-associated influencer having an opinion, and using that to choose a language.


> The language in question here is maintained by a BDFL, which means that one person has outsized influence on the language, and it's direction.

Seems like a strange perspective, isn't it expected (and even wanted) that if you have a BDFL, they have the most influence on the language and its direction? That's why they're BDFL, that cannot be outsized, it's perfectly sized for what it's explicitly trying to be, that's the entire point.

> In this context, I find it reasonable that if someone is ticked off by that BDFL, they might second guess the direction of the language itself. Since the opinions and emotions of that BDFL _will_ end up in the language and it's community.

Yeah, I suppose in the new daily reality where we care more about what opinions people hold, rather than what quality, reliability and so on they actually produce, this does make a lot of sense. Personally I never used Linux because of how much I love/despise Torvalds, I've basically never cared about what he thinks, never thought it was important either, as long as Linux continues to work for me.


Calling it an IDE is under-representing cursor

They have in-house models, and the data to train even more powerful ones. The cursor team is a proper AI lab.


> Calling it an IDE is under-representing cursor

On the contrary, it's over selling it: it's a not even a stand-alone IDE (like Zed, for instance) it's a mere fork of VSCode.


In the same sense that chrome is just a safari fork I suppose


Safari isn't even open source, and Chrome has never been a fork of Safari.

But yes Blink definitely started as a Webkit fork, and everyone would have found that laughable if someone bought that a proprietary fork of Webkit for $60B.


Isn't their in-house model just Kimi?


See here https://cursor.com/blog/composer-2-5

85% of the compute for the final model is from them, and not the base Kimi model.


That just means it cost a lot.

Does it perform meaningfully better than the Kimi model given all that extra compute? And proportionally to the amount spent?


Cursor's Composer 2.5 is one of the few models out there focusing on coding, which is the one thing most of us here want. It's pretty good! It's not near frontier level insight generating genius, but it's regarded as very capable and trustable, and is indeed a lot better than previous Kimi. It'll be interesting to compare it versus Kimi 2.7 Code, which just dropped, which is also notably a coding specifical model. I'm expecting we'll see more of this over time and I think it has huge rewards, and Composer 2.5 is early proof.

I'm not super concerned about the spend to train the model, especially given that Kimi was famously incredibly cheaply made, and given what they are competing with. I don't think that's a meaningful concern.

Reciprocally, and in far more important relevant in my humble opinion: in terms of cost to run models: Composer 2.5 is easily one of the cheapest models out there. It's fantastically cheap. It's token efficiency is through the roof astronomical. I think this training for a coding specific model has yielded something incredibly special here, and I hope SpaceXLAIC isn't the only company doing this.


That's something for us and benchmarks to decide

However it definitly isn't _just_ Kimi. The weight will be different after that 85% of extra training on top of the base model.

If those different weights are better are worse doesn't change that it's in most meaningful ways not the same as the base one.

I would encourage you to lookup their blog posts about their post training process if you want a bit more faith that they aren't running an extra 85% of compute and burning money with no-ops.


"Just Kimi" is hyperbole, to be clear.

I don't think it's all no-ops. Still don't think it's a particularly relevant model/company/product.

I'll defer the reading until I see signal that they have something worthwhile. I've watched a couple interviews and used the product, neither of which impressed me.


You don't think that a $60b valuation is having something worthwhile?

(Only half-joking…)


They use Kimi and post-train it on the same stuff that anyone with a Github dump can feed it. They aren't doing anything that you can't do yourself.


Dumping github into a model is not post training, thats pre training. And every base model already has all of github.

Composer post training is clearly very good, only second to Anthropic and OpenAI.

It does irk me a bit that they try to hide the fact that it's based on a chinese pretrained model though.


why comment on something you clearly don't know anything about? it's on-policy RL trained not just on coding text

listen and learn :)


Meh. On an outcomes analysis, I've found Cursor's delivery to be exceptionally weak.

Good luck to the alt-economy of SpaceTesla though, may all our 401ks survive.


Their “in house models” are reportedly basically just Kimi.


Replied on the other comment about this, but putting it here:

> See here https://cursor.com/blog/composer-2-5

> 85% of the compute for the final model is from them, and not the base Kimi model.

Of course they could be lying, but it seems feasible that they are adding a lot on top of this


I honestly much prefer this to the old way where the only mode of communication was speech or text. I now often understand a lot more holistically what the person coming with their product wants with just a demo + a conversation.

Of course you need the person making that vibe product to understand it's just a mockup of their idea and that it'll change. But I would argue this has always been a necessary quality for a product person.


thats really sad that only way you can understand something is with pictures and demos. like a little toddler.

when they already come up with a working model it doesnt really leave any room for abstract thought because now you are in concerte world . They see you as a machine that turns mocks into impls ( maybe you youself see yourself like that).

its ok if you see yourself as a code monkey monkey doing menial job of implementing someones mockup. But that job wont be here long and you will be on the streets holding "can turn working mocks to production code for food" signs.

honestly your attitude scares me more than some business guy building crap in lovable.


I can understand words, but having more diverse medias for communication lets a person express strictly more than before.

Sometimes words are better, sometimes a visual demo is better.

Is your solution to the problem you presented that you should artificially restrict what a person can express just to keep your own personal moat?

I prefer the alternative, let a person express themselves and grow thanks to AI, while keeping the necessary culture and boundaries to where it's also accepted for _me_ to cross boundaries and express my ideas to them in the same way. Or suggest other ways to express those ideas.

We then become a marketplace of ideas in a much deeper sense than before, where product managers would already expect you to implement what they want, but without them being able to convey it properly.

If I didn't have original ideas and didn't think I could compete in that marketplace of ideas, I would be scared like you convey in your message. But I'm confident that my value is not about translating things into code, it's because I have original thoughts I can convey to other people (and to AIs). (and about understanding architecture and systems to a degree that keeps me valuable even if all the code itself is written by AI without my direct involvement)


A (good) picture is worth 1024 words.


no business person will ever want to consider the right tradeoffs without being forced to

the ai making assumptions can’t get it right


Holy projection batman.

Might I suggest taking a step back and asking yourself why you were so triggered by the prior post. You made a bunch of mental leaps that are not supported by the prior post. Won't waste my time going through all of them, but the suggestion that they couldn't understand before has no basis in reality, they merely said it's easier to understand with increased fidelity of mocks.

The words can't hurt you. Deep breaths. :)

If your goal was discourse, might I suggest leaving the insults out of the message, it rarely is effectual.

If your goal was simply to insult, might I suggest leaving HNs, your anger and insults do the community a disservice.


What an insane replay.


I'm guessing the parent is wondering why this is noteworthy enough to be posted and discussed in this thread, and so if there's context they are missing


I'm wondering the same: why is this change significant enough to reach the frontpage on Hacker News?


Someone felt that was interesting and submitted it. People noticed Zig and upvoted. That would be my guess


Even before AI, deterministic checks by compilers are almost always better than "review the code"

"review the code" as a solution will eventually fail and cause a problem, even pre-AI.


The entire point of unsafe blocks and SAFETY comments is that they are easy for humans to find and audit, but not compiler checkable. If it can be compiler-checked by some clever token system, then ... it's just plain safe rust, and you don't need to document any special safety invariants in the first place


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: