> I have not been doing increasingly complex things since Opus 4.6 when models got really good.
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
I wonder if I would do better as a human, maybe? Or would I happen to have one move in 1500+ that's not valid?
I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece. Even through I do know the rules of chess, and I have played a few games once every so often.
It's always interesting since my train of thought always goes down:
- Ya they probably don't feel / think / have whatever living thing quality, they're just numbers on a machine going through calculations
- Wait, but am I not kind of the same thing? What is feeling for me if not basically the same thing?
- I have no clue if they think or feel or ....
Which in itself is a tired trope, but I also feel uncomfortable saying "These will never think / feel / ..." as an absolute
Regardless of that, I'm still going to interact with them, because even if they did feel, it would be in a way completely incomprehensible to us. There's not much point for me to try and cater to it's feelings at this point if that's the case. Nor is it possible in todays world to just avoid anything that is numbers being executed on a type of processor in case _everything_ has feelings.
You’re not just the same thing though. It’s sad that we’ve collectively forgotten that as we’ve gotten a better handle on the implementation details of the universe. That’s all science is, reverse engineering an inherent and eternal mystery.
I agree we aren't the same thing, but I would be curious for you to explain how you know with certainty why we don't share enough that we can rule out thinking / feeling / ... as things a model conceptually could do.
>I agree we aren't the same thing, but I would be curious for you to explain how you know with certainty why we don't share enough that we can rule out thinking / feeling / ... with certainty.
Stop anthropomorphizing these models. I understand it, we only have simple monkey brains to reason with and we can't help ourselves but draw comparisons to other things we see in nature. But these things are not alive.
I just want to note how you expressed your own feeling about the subject "these things are not alive", but without pointing to anything concrete explaining the theory you have behind this
And like I'm sure I'd agree depending on the definition of "alive" but then I'm also sure I would disagree depending on other definitions of "alive".
Hell, in biology alive is a hell of a topic these days. The grey space between dead and alive is much weirder than we ever expected. Really just points out how we're a persistent chemical reaction.
It's not that you're a d*, it's just that you lack any kind of nuance
There's a whole spectrum between self-hosting open weight models and having a cloud provider swap models from under you
Should you self host a model if want to maximize predictability to the limit? Yes. Does that mean it's wrong for someone hitting a model on API to expect that it won't switch to a completely different model under the hood from one day to another? Probably not.
Sure, but if you're working in a field where that limitation is not "because you like it" but "because you have to" it doesn't matter. If you cannot assume it to be true, then you have to assume it isn't.
Having experience breaking into networks is not the same thing as learning about the techniques used and the classes of vulnerabilities exploited by attackers.
As a guy who presumably has a lot less experience in security than you, I feel rude even bringing it up: surely you're aware of red teaming? This isn't a novel technique invented for AI- IBM has a page about it, it's what all the best DEF CON talks are about, it's the opening scene of Sneakers, it's the point of CtF games.
The language in question here is maintained by a BDFL, which means that one person has outsized influence on the language, and it's direction.
In this context, I find it reasonable that if someone is ticked off by that BDFL, they might second guess the direction of the language itself. Since the opinions and emotions of that BDFL _will_ end up in the language and it's community.
This is different than some un-associated influencer having an opinion, and using that to choose a language.
> The language in question here is maintained by a BDFL, which means that one person has outsized influence on the language, and it's direction.
Seems like a strange perspective, isn't it expected (and even wanted) that if you have a BDFL, they have the most influence on the language and its direction? That's why they're BDFL, that cannot be outsized, it's perfectly sized for what it's explicitly trying to be, that's the entire point.
> In this context, I find it reasonable that if someone is ticked off by that BDFL, they might second guess the direction of the language itself. Since the opinions and emotions of that BDFL _will_ end up in the language and it's community.
Yeah, I suppose in the new daily reality where we care more about what opinions people hold, rather than what quality, reliability and so on they actually produce, this does make a lot of sense. Personally I never used Linux because of how much I love/despise Torvalds, I've basically never cared about what he thinks, never thought it was important either, as long as Linux continues to work for me.
Safari isn't even open source, and Chrome has never been a fork of Safari.
But yes Blink definitely started as a Webkit fork, and everyone would have found that laughable if someone bought that a proprietary fork of Webkit for $60B.
Cursor's Composer 2.5 is one of the few models out there focusing on coding, which is the one thing most of us here want. It's pretty good! It's not near frontier level insight generating genius, but it's regarded as very capable and trustable, and is indeed a lot better than previous Kimi. It'll be interesting to compare it versus Kimi 2.7 Code, which just dropped, which is also notably a coding specifical model. I'm expecting we'll see more of this over time and I think it has huge rewards, and Composer 2.5 is early proof.
I'm not super concerned about the spend to train the model, especially given that Kimi was famously incredibly cheaply made, and given what they are competing with. I don't think that's a meaningful concern.
Reciprocally, and in far more important relevant in my humble opinion: in terms of cost to run models: Composer 2.5 is easily one of the cheapest models out there. It's fantastically cheap. It's token efficiency is through the roof astronomical. I think this training for a coding specific model has yielded something incredibly special here, and I hope SpaceXLAIC isn't the only company doing this.
However it definitly isn't _just_ Kimi. The weight will be different after that 85% of extra training on top of the base model.
If those different weights are better are worse doesn't change that it's in most meaningful ways not the same as the base one.
I would encourage you to lookup their blog posts about their post training process if you want a bit more faith that they aren't running an extra 85% of compute and burning money with no-ops.
I don't think it's all no-ops. Still don't think it's a particularly relevant model/company/product.
I'll defer the reading until I see signal that they have something worthwhile. I've watched a couple interviews and used the product, neither of which impressed me.
I honestly much prefer this to the old way where the only mode of communication was speech or text. I now often understand a lot more holistically what the person coming with their product wants with just a demo + a conversation.
Of course you need the person making that vibe product to understand it's just a mockup of their idea and that it'll change. But I would argue this has always been a necessary quality for a product person.
thats really sad that only way you can understand something is with pictures and demos. like a little toddler.
when they already come up with a working model it doesnt really leave any room for abstract thought because now you are in concerte world . They see you as a machine that turns mocks into impls ( maybe you youself see yourself like that).
its ok if you see yourself as a code monkey monkey doing menial job of implementing someones mockup. But that job wont be here long and you will be on the streets holding "can turn working mocks to production code for food" signs.
honestly your attitude scares me more than some business guy building crap in lovable.
I can understand words, but having more diverse medias for communication lets a person express strictly more than before.
Sometimes words are better, sometimes a visual demo is better.
Is your solution to the problem you presented that you should artificially restrict what a person can express just to keep your own personal moat?
I prefer the alternative, let a person express themselves and grow thanks to AI, while keeping the necessary culture and boundaries to where it's also accepted for _me_ to cross boundaries and express my ideas to them in the same way. Or suggest other ways to express those ideas.
We then become a marketplace of ideas in a much deeper sense than before, where product managers would already expect you to implement what they want, but without them being able to convey it properly.
If I didn't have original ideas and didn't think I could compete in that marketplace of ideas, I would be scared like you convey in your message. But I'm confident that my value is not about translating things into code, it's because I have original thoughts I can convey to other people (and to AIs). (and about understanding architecture and systems to a degree that keeps me valuable even if all the code itself is written by AI without my direct involvement)
Might I suggest taking a step back and asking yourself why you were so triggered by the prior post. You made a bunch of mental leaps that are not supported by the prior post. Won't waste my time going through all of them, but the suggestion that they couldn't understand before has no basis in reality, they merely said it's easier to understand with increased fidelity of mocks.
The words can't hurt you. Deep breaths. :)
If your goal was discourse, might I suggest leaving the insults out of the message, it rarely is effectual.
If your goal was simply to insult, might I suggest leaving HNs, your anger and insults do the community a disservice.
I'm guessing the parent is wondering why this is noteworthy enough to be posted and discussed in this thread, and so if there's context they are missing
The entire point of unsafe blocks and SAFETY comments is that they are easy for humans to find and audit, but not compiler checkable. If it can be compiler-checked by some clever token system, then ... it's just plain safe rust, and you don't need to document any special safety invariants in the first place
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
reply