Explore posts by Robin Ebers
2,249 posts by Robin Ebers
Select a post to readDeslopifying large AI codebases with subagentsinsightthe deslopification of larger AI projects (think: 100-200k lines of code, a month or two of work) will be studied
we can build at the speed of light, especially when you have UI/UX and engineering + architecture experience
but the time inevitably comes where you want to spend a few thousand dollars worth of tokens and run things like @cursor_ai's thermo-nuclear skill with 25 subagents on different parts of the app to remove 100k lines of slop
typical offenders:
- dead code (old/never removed/unreachable)
- overcomplification (too many branches)
- bandage-style fixes
- features you didn't even know existed
- outdated and unnecessary documentation
and honestly so much more
it's very satisfying when you learn how much code you can remove without the app changing it's behavior
here's an example prompt if you're token rich:
"run a massive deslopification workflow that aims to find DRY violations, remove dead code, overcomplicated code, clunky/fragile code, bandage fixes, or similar. first, learn our code and architecture, then look at leading open source projects in the space and compare our architecture to theirs. run as many subagents as you can. your north star is code simplification and massive removal of code while safely keeping the apps behavior materially the same"
you can also point them at specific open source projects that you know do something similar to you. for example, say "[..] open source projects like codex and opencode [...]" if you built AI agents
good luck!118200GPT-6 Astra is a faster Fable competitorguideGPT-6 Astra is here (and I tested it early)
beyond the obvious benchmark scores, GPT-6 Astra is absolutely insane in everyday real work
I've put over 21 billion tokens through this thing in a very short time, and i can honestly say:
THIS IS THE MODEL WE'VE BEEN WAITING FOR
a true and worthy Fable competitor (!)
GPT-6 Astra now can:
- design sites and apps that look incredible (see below)
- understand really complex problems better than before
- it writes significantly better than before (like, really good)
my favorite is probably that it doesn't talk to you like an autistic engineer anymore, but uses normal everyday language
BUT... it also is 2x the price of GPT 5.6 Sol (!), but in my tests it is not much cheaper than Fable 5.1, but also significantly faster
my brutally honest verdict:
1. if you can afford both, use both
2. if you can't, either one will now be good enough
3. if this model disappeared tomorrow, i'd be VERY sad
huge gg to the team
here's my unscientific design bench in al thinking levels
which one do you think did the best?
👉 https://t.co/QRNrh9ucmr 👈21195Ox Alpha edits clean video faster than OpusguideOx Alpha is a powerful model
i've put it into my new AI video editor
that editor currently runs on Fable 5 @ Low
and it... was fucking impressive
gave it two youtube videos (23 min and 28 min)
each video was edited 10 times
1) on the clean video it landed Opus 5 @ High's cut almost word for word (7 of 9 finished runs were 90–97% the same cut)
2) it edits for substance unprompted. deduped retakes down to the last good take, something Fable needs to be told to do.
3) ~7x faster than Opus 5 @ High on the clean video.
just a few problems...
- 4 of 20 runs broke: 1 invalid answer, 3 thought forever and never answered (all 3 on the harder video)
- it gets scattered on messy footage (runs agree 78% with each other there vs 88% on clean)
- we still don't know what it will cost
so one thing i can say for certain: on clean footage it edits at Opus 5 @ High level, which no model ever did. hard to ignore.
in a single sentence, on easier problems, the model outperforms Opus 5 @ High, but when things get messier, it falls behind fast
Ox Alpha is Opus's taste/insticts, but unfortunately without Opus's discipline.
BUT it is a lot faster!5322GPT 5.6 models find bugs Fable missesguideGPT 5.6 Sol/Terra/Luna is... fantastic!
some early/high-level observations from using it for a few hours
THE GOOD
1️⃣ it's indeed better at front. still not as good as Opus/Fable/GLM, but a measurable step forward and I'm here for it!
2️⃣ it immediately (!) found edge-cases/bugs in code that Fable wrote
3️⃣ 5.6 Terra especially got me excited because it is fast, cheap, and very capable
THE MEH
1️⃣ still talks very technical, even if asked not to
2️⃣ it consumed quite a bit of tokens, so 5x won't be enough
so I upgraded from 5x to 20x immediately to dig deeper
it's still early, but here are my 2 cents 👇🏻
GPT 5.6 Sol is very a very intelligent, diligent worker
it digs deeper than Fable 5, and finds things Fable couldn't
this is no longer a matter of "who wins" but "how to make them work together the best" and I'm REALLY excited to see where this is going
more to come!53880Claude Code visual previews simplify UI workinsightI fucking love Claude Code so much lately
and people keep asking me why
it's simple: because of moments like this 👇🏻
1. Claude Code gave me a huge plan
2. I obviously didn't want to read it
3. Instead just said: "tldr and preview"
Claude then used the visualize widget to preview what the UI changes in my app would look like
I don't know about you, but tI process things better when presented with them in a visual shape like this - it makes everything a million times easier3910Codex Remote connectivity remains slow and flakyguideCodex is great, except... whatever the underlying connectivity architecture is really fucking sucks
i genuinely don't understand how this can be so slow and so flaky
Codex Remote takes ages to load and update, controlling other devices from within Codex is laggy, sometimes it shows as disconnected or connections disappear
it's honestly been almost unusable
example: i'm currently on a plane with shit internet. i can talk to claude. i can talk to grok. but codex gives me all types of errors
- messages pretend to send but never do
- messages don't send but actually are sent
- "couldn't check worktree status" errors
- constant disconnects ("connection to x was lost...")
Cursor, Claude and Grok do this very well
their mobile apps are fucking great too
Codex' cloud infra is at least a year behind at this point
i'd literally pay extra to have a reliable cloud infra like Cursor or Devin
please fix, thank you554Blind tests reveal model differences in text editinginsightthis is gonna be very niche, but...
I just ran a lot more tests for TEXT EDITING (think: Cursor for video through transcripts)
the task is simple...
take a force-aligned transcript and edit it into a coherent social media post (simplified)
- Kimi K3
- Sonnet
- Opus 5
- Fable
- GPT 5.6 Sol
REALLY interesting results (if you care about how AI models think)
1️⃣ Kimi K3 Min/High/Max
good (the best non-Claude text model) but it has a lot of variance, meaning every edit came out vastly different
2️⃣ Sonnet Low/Medium/High/XHigh
surprisingly good, and I can see why Anthropic said this was such a great model. I'd never use it for coding, but for text even Low was 99% as good as High. XHigh was the closest to Opus 5 High. But it was much cheaper, and faster.
3️⃣ Opus 5 Low/High
Opus 5 is a good model, and in High it performs very well. but when you really look at its decisions over a handful of requests, you'll see patterns where it... just makes shit up. for example, when asked to cut a video, it set an arbitrary "40% target" which no one asked it to do. even worse: on Low effort it barely thought at all and the videos came out shit, while costing more, and taking longer. Even Sonnet 5 Low outperformed Opus 5 Medium.
4️⃣ Fable 5 Low
honestly, the show-stealer and a great example of why this model is so incredibly smart. not only was it much faster (about a third of the time) and much cheaper (about half of the cost), but despite being in Low effort, it outperformed every other variation by a landslide. for example, it didn't cut out a messy line early in the script because the payoff relied on it later.
Fable 5 is absolute king for 'thinking' through stuff, and doing so cheaper, faster and with more deterministic results
just for the lol's, I also threw in GPT 5.6 Sol (Medium and High) and while cheap and fast, it just couldn't compete with Fable. lots of variance (similar to Kimi K3) and no "big thinking" (see image below)
TL;DR:
- Sonnet is surprisingly good, even in low effort
- Opus is kinda bad, too expensive and too slow
- Fable 5 is an alien model that shouldn't exist
- GPT 5.6 Sol once again disappointed194Convex became Robin’s biggest workflow improvementinsightthere are few things in life that I appreciate more than having listened to @RayFernando1337 when it comes to Convex
he was definitely early but he was not wrong
@convex is genuinely the best quality-of-life improvement that my workflow has seen
i'd put it right after AI coding as a whole
that's how big of an impact it has made for me
i wish them nothing but enormous success373OpenUsage adds support for multiple Claude accountsguideOpenUsage v0.7.10-beta.3 now includes support for multiple Claude accounts
this is the first step to properly support them, in the way they were intended
here's what currently works:
- Claude Desktop and Claude Code
- iCloud sync across multiple devices
what's coming soon-ish:
- renaming of accounts
- other discoveries (home folders, cc-switch, etc)
- multiple Codex accounts
the current implementation is great for people who have e.g., a personal and a team account under the same email172Sonnet for most text, Fable for standout passagesguidemore text generation a/b blind testing
i stand by Sonnet being the best overall text model
it is incredibly good, very fast, very cheap, and very deterministic and reliable
it is definitely the one I'm going to start using in my apps now to generate most text
but the funny thing is that the second you add a little extra point for outstanding passages of text, Fable pulls ahead quite a lot
out of 10 text examples generated, 6 Fable generations (!) without knowing it, i marked as stand out good
that model is just a different beast
but it is also important to understand that Sonnet 5 is 20% of the cost of Fable
so the trick is:
- Sonnet for almost everything
- Fable through tool calls called by Sonnet
sick90GPT 5.6 Sol and Fable make different tradeoffsinsightGPT 5.6 SOL 🌞 ONE DAY LATER
the model generally is incredible. the limits are much higher. it's faster than fable 5 max. it is more diligent in its tasks. it can quite literally run for many hours and work autonomously, just like fable 5 would. terra is very useable too. better design.
what i don't like is this:
it has improved design performance, but honestly still lacks behind quite a bit. it burns tokens very fast for me and some others. the model is still an engineer, not your business partner, and it often talks a lot. it still can't write well. it still uses language that requires a lot of focus.
my two cents before the public release later today:
GPT 5.6 is going to be the best choice for engineers, and people needing the highest limits.
Fable 5 is still the best overall intelligence model, that is more capable at more things. writing. designing. overall thinking and ideas.
tl;dr: i'm going to use both. a lot.28655Codex and Fable helped build a playable gameinsightI handed this off to Codex two days ago
now using Fable 5 High/Max for adversarial reviews instead
it's been grinding and grinding and I can now successfully level to level 9
the game is REALLY playable
already done:
- 4 races (orc, undead, hornkin, troll)
- quests/experience/spawn/loot/etc
- actual character scaling, with stats and gear
- full fight mechanics to parry/miss/crit/dodge/etc
next:
- mage and hunter class
- a lot bigger world (current cap is level 10)
latest demo below (2x speed)8232GPT 5.6 speed could challenge ClaudeinsightOpus 4.8 and Fable 5 are great
but GPT 5.6 running about 10-15x FASTER than them...
while getting HIGHER LIMITS...
and the occasional STACKABLE limit reset...
even the most hardcore Claude fans might consider switching
(spoiler alert: I would)
because even if you need a few more prompts, the partnership with @cerebras will make it so much faster that you can iterate and steer the model to get better design done
that is, if the gracious U.S. makes the models available to us
another big thing that most people missed is this:
the new GPT 5.6 Terra (previously: mini) is about as good as Fable (in benchmarks), probably faster, and only $2.50 / $15 per million tokens
it's not a secret that I criticize OpenAI and the GPT-5 model series
but holy guacamole
THIS IS BEYOND A NEW PARADIGM
and it's exactly why I keep saying:
don't pick teams, use what is the best
and what's "the best" still changes quickly14515Robin maps equivalent levels across modelsguidekinda of agree with Tibo here, but not 100%
LIGHT is a great model that's quite powerful, but claiming Sol level is a stretch
MEDIUM is very good, daily driver
HIGH is the sweet spot for "i don't want to think about this" and "i don't want to spend 7 hours on this"
personally i'm mostly on MEDIUM and HIGH
while previously i was mostly on EXTRA HIGH
for absolute non-sense "explore and burn limits" i do run MAX sometimes and it's been great - but it also overengineers like crazy despite explicitly agreeing not to
TL;DR:
Tibo says two levels, i think it's closer to 1 level
light = medium, medium = high, etc
the same is true with Fable 5.1 btw5614Building a $2,000-a-month AI productinsightif you're wondering what i've been working on, i've spent the past 6 weeks building the most complex product i've ever built
it now sits at around 200k lines of code and has used about $50k worth of ai credits between codex, claude, and cursor
this started after an event in thailand with 200 other coaches that pretty much changed my life
shoutout to @takimoore btw
i've been using the most advanced ai models the world has ever seen and this thing has really shown me the limits of where we are today
and i've never been more convinced that coding is solved
the app now makes $2k a month and it's only the first month btw
there are 3 things i want to share here
1️⃣ coding is definitely solved, this product is only going to get bigger and the things i've learned building it have never stopped surprising me
2️⃣ when you put in enough work, enough passion, enough of your own experience and taste, nobody can compete with youfind that one specific problem you can solve in your own way, because if you just copy others then whatever you build is going to be easy to copy too
3️⃣ what you're seeing here is actually version 2, version 1 was a bunch of ai skills running in claude code with a free pilot groupthey validated the product and gave me the chance to build it this goodso never ever lock yourself in a room for months to build something nobody wants
if you're stuck, here's what to do:
figure out the ONE BIG PROBLEM you want to solve, validate it fast, and then build like your life depends on it
stop overthinking this shit
also, there are almost no coaches on twitter and i'm keeping this strictly for coaches for the time being, but if you are a coach feel free to dm me
new spots open next month205guidebet you didn't know this 👇🏻
when you run a Claude Code workflow in Ultracode and you run close to your limits, you can ask Claude to rewrite the ongoing workflow to gracefully stop
this can save a run that would otherwise fail, which is awesome https://t.co/MYhoougqMT94Seven deliberate reasons clients hate working with meguide7 reasons my clients hate working with me
every single one is deliberate
the ones who stay anyway get results. funny how that works lol https://t.co/48tasFlt53111guideanother day on AI twitter
another bunch of people complaining about the evil Anthropic
this isn't about Anthropic
it's about you not getting shit for free
all I see on my timeline is people crying that they won't have Fable 5 included in their subscription
but you whining about it won't change anything
you have two options:
1. use what you are given and be happy with it; or
2. vote with your wallet and move to something like a Codex
but for fuck sake, stop whining like a little bitch331Why I don't review the chinese modelsguidewhy I don't review the chinese models
let me reframe how I actually pick a model
1. what's this build worth to my business
2. which model gets it done best, not cheapest
3. use that one and ignore the fucking price
if a model would save me $1k/mo but cost me 10 hours more to make it work, that would value my time at $100/hour
I rarely do 1:1 calls but when I do, the investment is $1.5k/hour
everyone values their time differently
but for me the math just isn't mathing61guidei stopped using sol entirely
unusable for what i do unfortunately
in a nutshell:
- sol is a problem-creator (sees the world negatively)
- fable is a problem-solver (more realistic/positive)
grok def fits somewhere in-between
it sucks at writing almost as much as sol
but it's extremely fast, and great value
it's not an overall better model than sol
but it gets close to perfect for everyday stuff i'd argue20guideRT @skcd42: Now that Grok 4.5 is out, here are some of my workflows which I use daily
1. "use as many subagents and tokens as you need"
I…00tip@GregKara6 wild. i just don't feel any different.
the model does incredible work
- thousands of lines of PRs with no bugs/feedback
- analyzing UX in ways no other model ever could
- reading video files and transcripts and creating cowork projects that change my business
it's gold00GLM 5.2 changes the daily model choicetipswitched to GLM 5.2 approx. ~36 hours ago
unlike everywhere else, in Cursor this model KILLS
they even added image support for it which is wild
Cursor partnered with @FireworksAI_HQ on this
that means: zero data retention and very fast (80-100tps)
$145 (455M tokens) later and I don't miss Opus 4.8/GPT-5.5 yet
but all I know is this:
if all other models disappeared tomorrow, we'd be fine
only in the right harness though21850Cursor missing Astra limits its potentialinsightCursor not getting Astra is bad
like… really bad
it would likely have been the most powerful harness to get the most out of it
no doubt Grok 4.7 will be yet another big leap but there is zero chance it will be a Fable/Astra-level model
which is a bummer because i’m really warming up to Grok Bot, where at least we can orchestrate to Codex CLI
currently doing this on shit plane wifi
group chats are fucking wild10734Model efficiency cannot offset rising pricesinsightam a big fanboy for Astra/Codex right now but anyone claiming that token efficiency makes up for continuously DOUBLING PRICES is either delusional or lying to you
this has been a trend since GPT 5:
- GPT‑5: $1.25 in / $10 out
- GPT‑5.2: +40% in / +40% out
- GPT‑5.4: +43% in / +7% out
- GPT‑5.5: +100% in / +100% out
- GPT‑6 Astra: +100% in / +67% out
Astra is now $10 in / $50 out
that’s 8x on input and 5x on output
no token efficiency in the world makes up for this
and let’s not start talking about the priority multiplier that quietly went from 1.5x to 2x to a whopping 2.5x and recently back down to 2x
OpenAI is in no way shape or form “cheap”
not saying this to dunk on them
quite the opposite
this only shows that they’re confident that they’ve positioned themselves where they don’t need to discount anymore28327Fable 5’s re-release changes model accesstipON FABLE 5 RE-RELEASE
Thariq clarified that
1) the new classifiers shouldn't happen as often as the blog post might have suggested
2) it will be fully transparent when your request was rerouted (in logs I guess?)
3) you will NOT be charged Fable rates when a request was rerouted to Opus
this is a 10/10 update and how I wish Anthropic was always communicating
GG!11519AI changes what software teams can buildinsighti know this sounds crazy
because GPT 5.6 Sol is a great model
but if i had to choose between 5.6 Sol and Opus 4.8
i’d still pick Opus
hear me out on this
i promise this is not bait
claude models remain the best at writing, design, and “jack-of-all-trades” intelligence. they’re still the best at everyday tasks and obviously OpenAI knows this
GPT 6 is now rumored to be rushed to the finish line to better compete with Anthropic for probably that exact reason
if all models would disappear overnight and all i had was Sol, i’d be extremely happy. but the model isn’t quite the jump that Fable was, and it still lacks behind across the board
explicit examples include model behaviors too
GPT 5.x series models have this weird habit of digging themselves into a hole that is very hard to convince them to get out of
the only model doing this even more is Grok
Opus and Fable feel more natural, more agreeable, more eager to understand YOU and think about different angles
and despite having a $200 Codex sub, i still spend my time running down Fable in Claude, and switching to Opus for less important tasks instead of using GPT 5.6
it’s been a few days now and i’ve tested them A LOT
i’ll obviously continue to test and use them, and i’m impressed what OpenAI was able to squeeze out of the old model base
but it’s time to move on9918Cursor Sand can triage Slack, Sentry, and Linearguideso this seems to be what was previously rumored to be "Cursor Sand"
testing this, still early, but one thing I just set up:
7am every morning
- check slack channel for client feedback
- check sentry for new issues
- check linear for issues
- triage between all three
- validate in code
- propose top fixes
- delegate to cursor cloud
pretty wild4113Grok 4.6 and Fable make a strong pairinsightGrok 4.6 + Fable 5 is all you need
these two together are a killer combo
4.6 is especially great because it is less eager to finish the task. it can actually run for a while and check its own work.
if 4.7 truly is that much better (as Elon claims), i’d say OpenAI is in big trouble
Anthropic is fine, their ecosystem is much stronger and can absorb more pricing and model volatility
really curious about Fable 5.1 though718Codex can work autonomously for long sessionstipCodex is really doing this?
3 days into letting Codex build a local World of Warcraft clone
it's pretty insane so far
frame rate is poor, but it's been working for a few hours to fix that
🤣 https://t.co/wUk3FSHZUs375
No posts found
Try a different phrase or clear your filters.
Couldn't load the archive
Check your connection and try again.