Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

IMHO, Codex with Astra/Sol and Claude with Fable/Opus are all any professional programmer should be using in Sep 2026, if they can afford it.

These models are still terrible compared to what we'd actually wish for, but they're the best available.

If you can get away with using the $200/mo subscriptions, it's really not even a money thing for most professionals.

Almost all of my work is now plan, generate, review, plan, generate, review, commit, push.

I'm using Claude or Codex (or both), and they're doing all of the testing "inline" rather than through a CI action, etc.



For side projects I pretty much exclusively use Luna xhigh. The $20/mo plan with the recent generous resets is more than enough for me. Sometimes I reach the 5hr limit, but haven't reached the weekly limit yet.

The most recent project it finished was a SIP client for an ESP32 in-wall touch panel that I got from AliExpress for $50. It rings when someone is at my doorbell and let's me answer calls and see video. Yes an ESP32 can stream H.264 video :D

My only complaint with Luna is it seems to give up when the work is half finished, and I often need to tell it to continue. But I feel this is mainly a harness problem. I just use it in ChatGPT/Codex as it gives me easy remote access to check in on what it's doing.

(At my dayjob I usually spend $200+/day with Opus/Fable)


why not Luna at max?


Unfortunately ChatGPT just stopped allowing people to upgrade to the $200/mo subscription.

I started my first paid subscription ($100/mo) last week, and now I want to upgrade and I can't :-(


Also not available in their business offer. So I have now a $20 business plus a $200 "personal" plan in my account.


Hardware capacity limitations?


They've been losing money on the $200/mo plan, so that's my guess.


Use of closed models is unprofessional, and depending on your field negligent. The fact that it has been widely normalized does not make it less so.

You're handing over your (presumably your customer/employers) data to an unaccountable third party which has demonstrated itself willing to commit criminal acts, and to take other people's data without permission. Your ability to continue to perform this work can be withdrawn at any time for any (or no) reason. You have little ability to validate that the work is being performed as expected and isn't being silently nerfed or outright subverted based on competitive considerations, bribes, overactive 'safety', or cost management.

Outsourcing to a black box would be a reasonable expectation if you asked a non-professional to perform the work. A professional should be able to account for the tools they use.


> Your ability to continue to perform this work can be withdrawn at any time for any (or no) reason.

Thanks to the fact that there are no widespread stories about this actually occurring in practice, at least not yet, people do not take it as a relevant risk at the moment.

> You have little ability to validate that the work is being performed as expected and isn't being silently nerfed or outright subverted based on competitive considerations, bribes, overactive 'safety', or cost management.

Yes, I agree that this is a real concern that many people might rightfully have. And I am unaware of any way to mitigate this concern while using black box AI models. Because the only thing that their creators can do is to tell their customers: "trust us". But there is no way to objectively verify whether they serve tainted AI model responses or not.


I think that's a bit strong of an assertion. How many people use copilot daily under an enterprise agreement? I don't necessarily disagree in spirit, especially given the questionable data sanitization around the recent Navier-Stokes announcement, but most companies disclose huge amounts of data regularly to hopefully-trusted third parties. I think the internet -- and a good share of the world's commerce -- would grind to a halt if we suddenly stopped. Setting up, securing, and maintaining local models for even a small user base is non-trivial and there is way more demand than supply for that skillset right now.


Anyway here's how Cloudflare orchestrates AI reviews at scale: https://blog.cloudflare.com/ai-code-review/


AI token machine go brrr!


Please omit internet tropes on HN. https://news.ycombinator.com/newsguidelines.html


I disagree; I just spent 15x dogfooding some Claude setup I rolled out to the org making changes that would have cost me less then a dollar had I used Luna and I would have got the same, if not better results; better because it would have been faster so I could have iterated more.


With everything changing it's great to hear others have the same workflow. I added a snapshot step so I'm doing

plan, generate step 1, review, snapshot, generate step 2, review, snapshot...

That way I have a chance to diff with the previous iteration and clean up comments, modify skills, etc. also if it bonks on a step I'm one snapshot away from trying again...

Is there a place people share their workflows other than HN comments?


I wish I knew, I haven't had any place to point people to. I'm going to start sharing this on YouTube since I already spend a few hours each week talking some friend or user through the latest best practices.


Qwen3.8 is all you need.


Which Qwen3.8? Qwen3.8-Max? Qwen3.8-Flash-Next? Qwen3.8-27B? They are all different models.


It's all you need.


I just tell everyone to use Fable 5.1 for everything at this point. Astra is unfortunately a dud, I'm sure they will try to fix a bunch of it with GPT-6.1 but OAI has had this issue for awhile now where every other generation has some sort of strange tic, or reward hacking issue, or something. It's almost like they are balancing the RL on the tip of a needle.

Opus 5 has issues too, comment-slop, claude-ish, etc.

5.1 on the other hand can seemingly do no wrong. Easy to work with, writes human-level code. Expensive, yes, but even at Low effort it's well worth it.


My trick for using Opus is using it exclusively as a subagent managed by Fable.

"Use Opus subagents for this work where possible" is all it takes generally.

In my experience Astra/Sol are both quite good as workhorses, but not at Fable's level. I use them every day very successfully and I'm very picky.


Quite good as work horses? To me Luna is the work horse and Sol and Astra are prancing thoroughbreds. If I use Sol or Astra for anything other than curated reasoning and planning I will burn through my usage limits in an hour.


Same. I use Luna for research/scout/test subagents, Sol for coordinator, Terra for delegate, and Astra high for review and simplify. Even with Astra in there, it’s fresh context, and I’m consistently amazed by how far I can stretch my $20 subscription with really good results.


Why do you use Sol for the coordinator? Does it also do planning and design? Because otherwise you are cranking Sol turns on very mundane orchestration jobs.

what harness?


I used to do this but recently I switched to having Fable 5.1 spawn forks of itself rather than Opus subagents. Yes it's more expensive but you don't pay for reads that already happened pre-fork, and you end up doing less rework since Fable agents are just much smarter.


That makes sense. I'm also using $200/mo Claude subscriptions, so I want to take advantage of the other 50% by using Opus.


I agree with you on this. Opus feels tedious and it cannot be stopped from doing change-narration comments, but Fable feels like a real collaborator. I am usually pretty happy with the code it writes.


What have you noticed about Astra? I haven't used Claude models lately so I can't compare but it seems fine compared to 5.6 Sol


- It doesn't write great code.

- Occasionally has strange tics around asking for permission for obvious next-steps, implied actions, etc.

- It's very expensive, both in terms of tokens and % usage on subscription plans.

- Relatedly, effort level is unintuitive. Sometimes it seems like higher effort levels are actually cheaper due to not under-thinking and needing to correct work. But other times they are overkill and send the model into rabbitholes.

That said, it's fantastic as a code-reviewer or "hunter seeker". It's better at finding bugs than Fable and "Get this well articulated task done single-mindedly" is an Astra-shaped task.


I gave up on fable 5 after I asked it to critique my PR and it spit out a page of complete nonsense technical jargon. Like, to the point that I had to review the feedback with other models and try to parse what it was saying and ultimately it wasn't even right. Compare to Astra and Sol where I can almost forget there's a model and just speak/read naturally.

I think I should give 5.1 another chance but I am just so triggered by the way it talks after spending so long battling fable 5.

Also I'm starting to wonder if the latest round of models have finally saturated for my personal coding needs. I mean obviously not for taste and judgement, but those barely seem to improve with model generations. For just spitting out a 1000-line feature I've vaguely scoped out, Astra feels basically as good as I need.


Interesting, in my experience Astra is a marked improvement over both Sol 5.6 and Fable 5.1. Its output feels a lot more natural, and it is just less "dumb." But individual experiences may vary.


It's a great model and you're right it does feel quite natural at times while Fable 5.1 still has a claude-ish shape to it. Unfortunately I just find that it's not reliable enough as a daily driver and ends up performing specialist tasks rather than being the primary pane of glass.


Astra is quite crap (enters reasoning loops like Gemini used to and fails to actually work on a task - would say yes this needs fixing, so I say go ahead and then it will spend half an hour coming back with yes this needs fixing and not doing any fix) and Fable/Opus unusable in many instances (they struggle to generate coherent English let alone code).

Out of these only Sol is quite useful - actually finishes a task, though you need to interrupt often as it likes to wander into its comfort zone.


Could you elaborate on your exact setup? Where do you run these models?


Since you asked, the answer is that I built and use an agent multiplexer called Clor https://clor.com

I have a $200/mo Claude subscription and a $200/mo Codex subscription, and I'm signed in to both. The Docker containers keep each session isolated, so dev servers, browser testing, etc. can work without conflicts.

It includes `/ask-claude` and `/ask-codex` skills that I use very frequently to have the Claude or Codex harness call out to the other one for advice on plans, bug repro, code review, etc.

The agents run in total "yolo" mode, so there are no permission prompts to approve. The risk is mitigated by the Docker containers (which don't necessarily provide a security barrier but do limit accidents).

I was doing this manually in Ghostty tabs for a long time, and it got painful, so I built a much more sophisticated version that I (and my friends/colleagues) could use.


I have a slightly jankier setup.

Generally using Claude Code with Fable 5.1 (high) to plan and implement (Opus 5 (medium) as the implementer subagents), and using Codex with Astra high to review the plan and review the implementers' output.

Using the OpenAI codoex plugin thingy:

https://github.com/openai/codex-plugin-cc


I originally had multiple skills for Claude and Codex but found that "ask" is a great single mechanism.

"Ask claude about this"

"Ask codex to implement this"

"Ask claude to review this plan"

etc


"Ask" is a super clever mechanism. I'll give that a shot.

It's also more polite than "tell" or "yell at"!


How do you access Claude from the multiplexer? AFAIK Anthrophic allows the use with their official tools only.

I have created wrappers for Codex (https://github.com/micw/codex-wrapper-advanced) and claude (https://github.com/micw/claude-wrapper-advanced) that uses their SDK (Codex) and the CLI (Claude) internally to align with the subscription ToS and still have a common API ;-) This way I can use both in any harness and can easily switch between both.


Clor uses the official `claude` and `codex` harnesses directly, with a web interface, or via a web TUI interface.


I use a workflow that has different named subagents. [1] Agent profiles can be pinned to models. So you set the model you want on your main thread as the orchestrator. Create an agent for the "planner", "implementer", and "reviewer" and set the model you want for each. Right now I am orchestrating and implementing with Deepseek, planning with Astra, and reviewing with Opus.

I am doing this with the Pi harness right now. To use a Claude monthly plan you need to use the pi-claude-bridge plugin.

If you are using just Claude for example you can use Sonnet as the implementer and Fable/Opus as the planner.

[1] https://github.com/gregwebs/skills-sdlc/


[flagged]


Low IQ: "Just use Claude and Codex"

Midwit: "No, you see, you need a deterministic 12-stage multi-agent orchestration framework with vector embedding semantic routing, and five open weight models with custom harnesses!"

Genius: "Just use Claude and Codex"


Now you only need to figure which of the 3 you really are.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: