From my memory of the 80’s, the introverts dealt with the stress of dealing with so many people for so long by quietly getting drunk enough to talk to people.
I bought a 5090 a year an a half ago for $2000. The same card, now a year and a half older, is $4000. Then there is the RAM - I bought 96GB, wishing it was 128, and now the price on my old RAM has doubled.
> Then there is the RAM - I bought 96GB, wishing it was 128, and now the price on my old RAM has doubled.
I also bought 96GB some while ago but after the initial increases, thinking I'll wait it out. Now 128GB is far more expensive than it was when I first looked. Luck has it I want DDR5 RDIMM as well, which seems the hardest hit when it comes to RAM prices, fun stuff.
There is a lot of "AI demand" that isn't just running inference on an LLM whose weights you downloaded.
I'm training a model using reinforcement learning with self-play. I can and do use vast.ai when scaling but for experiments it's far faster, and cheaper, to run it locally until the bugs are all figured out. Just provisioning a new instance and copying the relevant checkpoints and things can take 25 minutes. It's zero locally.
Likewise. I have a huge demand personally to run AI noise-filtering models on many TB per month of raw video files. It takes about 3 days per file.
Apples ProRes codec is only licensed to run in high quality mode on a Mac, and so my Nvidia PC can’t do what I need. Thus, I own the beefiest Mac Studio you can currently buy. I would pay more for more TFlops.
I have done local LLM on there but it wasn’t interesting. Far worse performance and intelligence per dollar than the cloud boys.
There is no cloud offering for my video needs though.
It’s probably files that, over the course of a month, add up to multiple TBs.
Which would suggest that a 1Gbit fibre connection would be adequate. For serious commercial usage, multi Gbit fibre is available in many places around the world.
Although being video files, they could easily be in the TB range. In which case, it would be interesting to know!
10mbps constant = ~3TiB a month, which is close to what I assume they're doing to acquire the video (that's about what a few SD x264 streams would run). So you don't even need a fast connection, per se.
A C-level executive I know is getting a top-of-the-line new Mac simply to function as a personal build server and host for agentic coding instances - they are able to orchestrate so many parallel projects that they're hitting RAM limits from sessions and the builds and local test runs they're kicking off (largly unsupervised). Before AI, they'd only had a MacBook Air; this completely changes their workflows. They talk about how many other executives they've met are equally giddy at having gone from coding few to no projects themselves, to coding more projects in parallel than any of their respective pre-AI technical colleagues.
I'd suspect that agentic coding has birthed so many new effective engineers, that the entire dynamics of demand for high-end machines have been upended.
The general problem I have with this is that I think those C-level executive should be doing C-level stuff and not programmer stuff. If you're using AI to help with your main job, great! If you're using AI to do someone else's work, then I think maybe you should change jobs. If you want to be a programmer, go for it! But not at the expense of your "real" job.
I'm pretty sure the C-level people don't want us programmers faffing around with their roles. Why are they encroaching into ours? Don't they already have enough to do; setting the direction of the company, making sure it's profitable, selling things to people, etc.?
It very much depends what projects they are coding. One thing agentic coding is really good for is developing highly specific custom applications specifically for your own use case. This can include the kinds of decision support and analytics tasks a C-level relies on to understand how their organisation is functioning, what's actually going on, and where the gaps are.
My team is using coding agents to develop analytics and reporting applications, and an engineering automation application, specific to our needs, that simply would not exist otherwise. The cost and time would be far too great. I can easily imagine there are plenty of cases like this, for all sorts of roles, including managerial ones, where only the person doing the job understands the job well enough to specify a requirement and guide an agent to code exactly what they need.
ram? coding instances? i suspect this is not doing any inference on the machine and i also suspect the ram use is due to a bagilion node/python processes running.
nothing like 600mb of for type check server another 500 for webpack and another 600 for the inevitable electron wrapper you didnt know about. python isnt as bad but still not great.
I’ve been happy training and running inference for small language models on my M4 Mac.
Inference with MLX is surprisingly zippy. I’m running a classification task on the entire HN comment dataset and it’s projected to take about two and a half days, which is not bad considering we’re talking about tens of millions of comments.
Yes, I could do it much more quickly by throwing Modal GPUs at it but this is low-priority work. I might as well throw my M4 a bone.
Is Modal at all similar to Vast.ai or just related because "It's for AI"? I looked at Modal's page for training, and it talks about using some SDK and other junk, can you not just get a beefy instance from Modal with tons of VRAM to do what you want with?
Same, but with vision models. Unfortunately, I might be at my limit locally. I have three models that I'm using to find and identify objects in pictures. The largest dataset and model now takes about 8 hours per epoch on my Mac M4 with 16G memory.
Yeah this was what got me to start doing short rentals of bigger gpus in the clouds, upload your parquet files and it takes a couple of hours for a thing that would have my mac at 100% for a couple of days
Yes, basically like alphago. I’m teaching it to play magic: the gathering.
I had to start with some heuristic-based bots that played the decks very simply just to get to the point where the was some signal to learn from. I did behavioral cloning on the bots as a foundation, then self-play.
Do you find that CoreML manages to fill up your drive with so many tiny files that a reboot takes hours to clean them up? I keep meaning to get my friends still inside the spaceship to file a radar about that.
I feel like I'm the only person on HN that doesn't have any major (or cliche) problems with GCP, even after using it for a decade at this point. Like it's not perfect and I've hit weird roadblocks along the way but I've dabbled AWS and I currently use Azure at work and those are hilariously bad in their own ways, people just seem to kind of be used to it?
I like GCP as well, having previously used AWS for many years.
The biggest issue is that it's taken Google a long time to figure out what enterprise security is - it's not something what was in their DNA - and the product shows it. They're too focused on weird fancy half-assed "beyond" solutions to properly implement the basics. They dangle clever federated etc. solutions that only work in some circumstances so that you're better off just avoiding them, but the more basic alternatives are limited in their own way because they focus too much on the fancier solutions.
I find GCP's developer experience to be superior to AWS. Maybe it's just me, but I found Cloud Run to be super simple to use compared to ECS/Fargate. The "gcloud" command structure also makes more sense to me.
Thermal performance degrades too. Paste breaks down over 2 to 5 years, you'll thermal throttle more often and more quickly than when new after a few years.
Phase change thermal pads help with this, extending that timeline out beyond 5+ years but still will eventually degrade.
There's all kinds of other hardware things we don't usually think about that can start having an actual performance impact after ~3-5 years too. Battery, obviously, as it degrades can limit your peak wattage, the power delivery capacitors lose capacitance over the years and develop more resistance impacting maximum clock states, fan bearings wear out, you get vacuum leaks in copper heat pipes, and micro-fissures in soldering.
Granted, most of it is going to be software but its easy to forget that hardware degrades, and degrades relatively quickly especially in higher power systems.
This is why repairability and user serviceability is so important. If you can't open up and swap parts in your machine easily, a $10 consumable turns it into $2,000+ ewaste.
Because a month later is a month of recursive self improvement at the speed of light. Once a lab catches up, the first lab will be TWO months ahead, then a year ahead, then forever ahead. After a few months of this the differences in absolute terms will be enormous.
I’m still at the point where Fable is still very stupid and needs constant oversight and correction and questioning to keep it on task. Anything less would be close to unusable.
I did not know that - but what happens then I get denied and I plead with the model that my usecase doesn't violate its TOS ? Edit prompt till I get it to pass ?
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
Yes, you get to re-try it until it might pass. Each attempt uses tokens at Fable level though. I once burned through a 4 hour window ($20 plan) in 4 attempts to re-frame the request.
You want to disable "Switch models when a message is flagged"[1]
With less people, this happens less nowadays.
reply