xwisdom.

Explore posts by Robin Ebers

2,249 posts by Robin Ebers

Select a post to read
Deslopifying large AI codebases with subagentsinsightthe deslopification of larger AI projects (think: 100-200k lines of code, a month or two of work) will be studied we can build at the speed of light, especially when you have UI/UX and engineering + architecture experience but the time inevitably comes where you want to spend a few thousand dollars worth of tokens and run things like @cursor_ai's thermo-nuclear skill with 25 subagents on different parts of the app to remove 100k lines of slop typical offenders: - dead code (old/never removed/unreachable) - overcomplification (too many branches) - bandage-style fixes - features you didn't even know existed - outdated and unnecessary documentation and honestly so much more it's very satisfying when you learn how much code you can remove without the app changing it's behavior here's an example prompt if you're token rich: "run a massive deslopification workflow that aims to find DRY violations, remove dead code, overcomplicated code, clunky/fragile code, bandage fixes, or similar. first, learn our code and architecture, then look at leading open source projects in the space and compare our architecture to theirs. run as many subagents as you can. your north star is code simplification and massive removal of code while safely keeping the apps behavior materially the same" you can also point them at specific open source projects that you know do something similar to you. for example, say "[..] open source projects like codex and opencode [...]" if you built AI agents good luck!Robin EbersAug 26200GPT-6 Astra is a faster Fable competitorguideGPT-6 Astra is here (and I tested it early) beyond the obvious benchmark scores, GPT-6 Astra is absolutely insane in everyday real work I've put over 21 billion tokens through this thing in a very short time, and i can honestly say: THIS IS THE MODEL WE'VE BEEN WAITING FOR a true and worthy Fable competitor (!) GPT-6 Astra now can: - design sites and apps that look incredible (see below) - understand really complex problems better than before - it writes significantly better than before (like, really good) my favorite is probably that it doesn't talk to you like an autistic engineer anymore, but uses normal everyday language BUT... it also is 2x the price of GPT 5.6 Sol (!), but in my tests it is not much cheaper than Fable 5.1, but also significantly faster my brutally honest verdict: 1. if you can afford both, use both 2. if you can't, either one will now be good enough 3. if this model disappeared tomorrow, i'd be VERY sad huge gg to the team here's my unscientific design bench in al thinking levels which one do you think did the best? 👉 https://t.co/QRNrh9ucmr 👈Robin EbersSep 2695Ox Alpha edits clean video faster than OpusguideOx Alpha is a powerful model i've put it into my new AI video editor that editor currently runs on Fable 5 @ Low and it... was fucking impressive gave it two youtube videos (23 min and 28 min) each video was edited 10 times 1) on the clean video it landed Opus 5 @ High's cut almost word for word (7 of 9 finished runs were 90–97% the same cut) 2) it edits for substance unprompted. deduped retakes down to the last good take, something Fable needs to be told to do. 3) ~7x faster than Opus 5 @ High on the clean video. just a few problems... - 4 of 20 runs broke: 1 invalid answer, 3 thought forever and never answered (all 3 on the harder video) - it gets scattered on messy footage (runs agree 78% with each other there vs 88% on clean) - we still don't know what it will cost so one thing i can say for certain: on clean footage it edits at Opus 5 @ High level, which no model ever did. hard to ignore. in a single sentence, on easier problems, the model outperforms Opus 5 @ High, but when things get messier, it falls behind fast Ox Alpha is Opus's taste/insticts, but unfortunately without Opus's discipline. BUT it is a lot faster!Robin EbersAug 2622GPT 5.6 models find bugs Fable missesguideGPT 5.6 Sol/Terra/Luna is... fantastic! some early/high-level observations from using it for a few hours THE GOOD 1️⃣ it's indeed better at front. still not as good as Opus/Fable/GLM, but a measurable step forward and I'm here for it! 2️⃣ it immediately (!) found edge-cases/bugs in code that Fable wrote 3️⃣ 5.6 Terra especially got me excited because it is fast, cheap, and very capable THE MEH 1️⃣ still talks very technical, even if asked not to 2️⃣ it consumed quite a bit of tokens, so 5x won't be enough so I upgraded from 5x to 20x immediately to dig deeper it's still early, but here are my 2 cents 👇🏻 GPT 5.6 Sol is very a very intelligent, diligent worker it digs deeper than Fable 5, and finds things Fable couldn't this is no longer a matter of "who wins" but "how to make them work together the best" and I'm REALLY excited to see where this is going more to come!Robin EbersJul 2680Claude Code visual previews simplify UI workinsightI fucking love Claude Code so much lately and people keep asking me why it's simple: because of moments like this 👇🏻 1. Claude Code gave me a huge plan 2. I obviously didn't want to read it 3. Instead just said: "tldr and preview" Claude then used the visualize widget to preview what the UI changes in my app would look like I don't know about you, but tI process things better when presented with them in a visual shape like this - it makes everything a million times easierRobin EbersJun 2610Codex Remote connectivity remains slow and flakyguideCodex is great, except... whatever the underlying connectivity architecture is really fucking sucks i genuinely don't understand how this can be so slow and so flaky Codex Remote takes ages to load and update, controlling other devices from within Codex is laggy, sometimes it shows as disconnected or connections disappear it's honestly been almost unusable example: i'm currently on a plane with shit internet. i can talk to claude. i can talk to grok. but codex gives me all types of errors - messages pretend to send but never do - messages don't send but actually are sent - "couldn't check worktree status" errors - constant disconnects ("connection to x was lost...") Cursor, Claude and Grok do this very well their mobile apps are fucking great too Codex' cloud infra is at least a year behind at this point i'd literally pay extra to have a reliable cloud infra like Cursor or Devin please fix, thank youRobin EbersSep 264Blind tests reveal model differences in text editinginsightthis is gonna be very niche, but... I just ran a lot more tests for TEXT EDITING (think: Cursor for video through transcripts) the task is simple... take a force-aligned transcript and edit it into a coherent social media post (simplified) - Kimi K3 - Sonnet - Opus 5 - Fable - GPT 5.6 Sol REALLY interesting results (if you care about how AI models think) 1️⃣ Kimi K3 Min/High/Max good (the best non-Claude text model) but it has a lot of variance, meaning every edit came out vastly different 2️⃣ Sonnet Low/Medium/High/XHigh surprisingly good, and I can see why Anthropic said this was such a great model. I'd never use it for coding, but for text even Low was 99% as good as High. XHigh was the closest to Opus 5 High. But it was much cheaper, and faster. 3️⃣ Opus 5 Low/High Opus 5 is a good model, and in High it performs very well. but when you really look at its decisions over a handful of requests, you'll see patterns where it... just makes shit up. for example, when asked to cut a video, it set an arbitrary "40% target" which no one asked it to do. even worse: on Low effort it barely thought at all and the videos came out shit, while costing more, and taking longer. Even Sonnet 5 Low outperformed Opus 5 Medium. 4️⃣ Fable 5 Low honestly, the show-stealer and a great example of why this model is so incredibly smart. not only was it much faster (about a third of the time) and much cheaper (about half of the cost), but despite being in Low effort, it outperformed every other variation by a landslide. for example, it didn't cut out a messy line early in the script because the payoff relied on it later. Fable 5 is absolute king for 'thinking' through stuff, and doing so cheaper, faster and with more deterministic results just for the lol's, I also threw in GPT 5.6 Sol (Medium and High) and while cheap and fast, it just couldn't compete with Fable. lots of variance (similar to Kimi K3) and no "big thinking" (see image below) TL;DR: - Sonnet is surprisingly good, even in low effort - Opus is kinda bad, too expensive and too slow - Fable 5 is an alien model that shouldn't exist - GPT 5.6 Sol once again disappointedRobin EbersAug 264Convex became Robin’s biggest workflow improvementinsightthere are few things in life that I appreciate more than having listened to @RayFernando1337 when it comes to Convex he was definitely early but he was not wrong @convex is genuinely the best quality-of-life improvement that my workflow has seen i'd put it right after AI coding as a whole that's how big of an impact it has made for me i wish them nothing but enormous successRobin EbersSep 263OpenUsage adds support for multiple Claude accountsguideOpenUsage v0.7.10-beta.3 now includes support for multiple Claude accounts this is the first step to properly support them, in the way they were intended here's what currently works: - Claude Desktop and Claude Code - iCloud sync across multiple devices what's coming soon-ish: - renaming of accounts - other discoveries (home folders, cc-switch, etc) - multiple Codex accounts the current implementation is great for people who have e.g., a personal and a team account under the same emailRobin EbersAug 262Sonnet for most text, Fable for standout passagesguidemore text generation a/b blind testing i stand by Sonnet being the best overall text model it is incredibly good, very fast, very cheap, and very deterministic and reliable it is definitely the one I'm going to start using in my apps now to generate most text but the funny thing is that the second you add a little extra point for outstanding passages of text, Fable pulls ahead quite a lot out of 10 text examples generated, 6 Fable generations (!) without knowing it, i marked as stand out good that model is just a different beast but it is also important to understand that Sonnet 5 is 20% of the cost of Fable so the trick is: - Sonnet for almost everything - Fable through tool calls called by Sonnet sickRobin EbersAug 260GPT 5.6 Sol and Fable make different tradeoffsinsightGPT 5.6 SOL 🌞 ONE DAY LATER the model generally is incredible. the limits are much higher. it's faster than fable 5 max. it is more diligent in its tasks. it can quite literally run for many hours and work autonomously, just like fable 5 would. terra is very useable too. better design. what i don't like is this: it has improved design performance, but honestly still lacks behind quite a bit. it burns tokens very fast for me and some others. the model is still an engineer, not your business partner, and it often talks a lot. it still can't write well. it still uses language that requires a lot of focus. my two cents before the public release later today: GPT 5.6 is going to be the best choice for engineers, and people needing the highest limits. Fable 5 is still the best overall intelligence model, that is more capable at more things. writing. designing. overall thinking and ideas. tl;dr: i'm going to use both. a lot.Robin EbersJul 2655Codex and Fable helped build a playable gameinsightI handed this off to Codex two days ago now using Fable 5 High/Max for adversarial reviews instead it's been grinding and grinding and I can now successfully level to level 9 the game is REALLY playable already done: - 4 races (orc, undead, hornkin, troll) - quests/experience/spawn/loot/etc - actual character scaling, with stats and gear - full fight mechanics to parry/miss/crit/dodge/etc next: - mage and hunter class - a lot bigger world (current cap is level 10) latest demo below (2x speed)Robin EbersJul 2632GPT 5.6 speed could challenge ClaudeinsightOpus 4.8 and Fable 5 are great but GPT 5.6 running about 10-15x FASTER than them... while getting HIGHER LIMITS... and the occasional STACKABLE limit reset... even the most hardcore Claude fans might consider switching (spoiler alert: I would) because even if you need a few more prompts, the partnership with @cerebras will make it so much faster that you can iterate and steer the model to get better design done that is, if the gracious U.S. makes the models available to us another big thing that most people missed is this: the new GPT 5.6 Terra (previously: mini) is about as good as Fable (in benchmarks), probably faster, and only $2.50 / $15 per million tokens it's not a secret that I criticize OpenAI and the GPT-5 model series but holy guacamole THIS IS BEYOND A NEW PARADIGM and it's exactly why I keep saying: don't pick teams, use what is the best and what's "the best" still changes quicklyRobin EbersJun 2615Robin maps equivalent levels across modelsguidekinda of agree with Tibo here, but not 100% LIGHT is a great model that's quite powerful, but claiming Sol level is a stretch MEDIUM is very good, daily driver HIGH is the sweet spot for "i don't want to think about this" and "i don't want to spend 7 hours on this" personally i'm mostly on MEDIUM and HIGH while previously i was mostly on EXTRA HIGH for absolute non-sense "explore and burn limits" i do run MAX sometimes and it's been great - but it also overengineers like crazy despite explicitly agreeing not to TL;DR: Tibo says two levels, i think it's closer to 1 level light = medium, medium = high, etc the same is true with Fable 5.1 btwRobin EbersSep 2614Building a $2,000-a-month AI productinsightif you're wondering what i've been working on, i've spent the past 6 weeks building the most complex product i've ever built it now sits at around 200k lines of code and has used about $50k worth of ai credits between codex, claude, and cursor this started after an event in thailand with 200 other coaches that pretty much changed my life shoutout to @takimoore btw i've been using the most advanced ai models the world has ever seen and this thing has really shown me the limits of where we are today and i've never been more convinced that coding is solved the app now makes $2k a month and it's only the first month btw there are 3 things i want to share here 1️⃣ coding is definitely solved, this product is only going to get bigger and the things i've learned building it have never stopped surprising me 2️⃣ when you put in enough work, enough passion, enough of your own experience and taste, nobody can compete with youfind that one specific problem you can solve in your own way, because if you just copy others then whatever you build is going to be easy to copy too 3️⃣ what you're seeing here is actually version 2, version 1 was a bunch of ai skills running in claude code with a free pilot groupthey validated the product and gave me the chance to build it this goodso never ever lock yourself in a room for months to build something nobody wants if you're stuck, here's what to do: figure out the ONE BIG PROBLEM you want to solve, validate it fast, and then build like your life depends on it stop overthinking this shit also, there are almost no coaches on twitter and i'm keeping this strictly for coaches for the time being, but if you are a coach feel free to dm me new spots open next monthRobin EbersSep 265guidebet you didn't know this 👇🏻 when you run a Claude Code workflow in Ultracode and you run close to your limits, you can ask Claude to rewrite the ongoing workflow to gracefully stop this can save a run that would otherwise fail, which is awesome https://t.co/MYhoougqMTRobin EbersJul 264Seven deliberate reasons clients hate working with meguide7 reasons my clients hate working with me every single one is deliberate the ones who stay anyway get results. funny how that works lol https://t.co/48tasFlt53Robin EbersJul 261guideanother day on AI twitter another bunch of people complaining about the evil Anthropic this isn't about Anthropic it's about you not getting shit for free all I see on my timeline is people crying that they won't have Fable 5 included in their subscription but you whining about it won't change anything you have two options: 1. use what you are given and be happy with it; or 2. vote with your wallet and move to something like a Codex but for fuck sake, stop whining like a little bitchRobin EbersJul 261Why I don't review the chinese modelsguidewhy I don't review the chinese models let me reframe how I actually pick a model 1. what's this build worth to my business 2. which model gets it done best, not cheapest 3. use that one and ignore the fucking price if a model would save me $1k/mo but cost me 10 hours more to make it work, that would value my time at $100/hour I rarely do 1:1 calls but when I do, the investment is $1.5k/hour everyone values their time differently but for me the math just isn't mathingRobin EbersJun 261guidei stopped using sol entirely unusable for what i do unfortunately in a nutshell: - sol is a problem-creator (sees the world negatively) - fable is a problem-solver (more realistic/positive) grok def fits somewhere in-between it sucks at writing almost as much as sol but it's extremely fast, and great value it's not an overall better model than sol but it gets close to perfect for everyday stuff i'd argueRobin EbersAug 260guideRT @skcd42: Now that Grok 4.5 is out, here are some of my workflows which I use daily 1. "use as many subagents and tokens as you need" I…Robin EbersJul 260tip@GregKara6 wild. i just don't feel any different. the model does incredible work - thousands of lines of PRs with no bugs/feedback - analyzing UX in ways no other model ever could - reading video files and transcripts and creating cowork projects that change my business it's goldRobin EbersJul 260GLM 5.2 changes the daily model choicetipswitched to GLM 5.2 approx. ~36 hours ago unlike everywhere else, in Cursor this model KILLS they even added image support for it which is wild Cursor partnered with @FireworksAI_HQ on this that means: zero data retention and very fast (80-100tps) $145 (455M tokens) later and I don't miss Opus 4.8/GPT-5.5 yet but all I know is this: if all other models disappeared tomorrow, we'd be fine only in the right harness thoughRobin EbersJun 2650Cursor missing Astra limits its potentialinsightCursor not getting Astra is bad like… really bad it would likely have been the most powerful harness to get the most out of it no doubt Grok 4.7 will be yet another big leap but there is zero chance it will be a Fable/Astra-level model which is a bummer because i’m really warming up to Grok Bot, where at least we can orchestrate to Codex CLI currently doing this on shit plane wifi group chats are fucking wildRobin EbersSep 2634Model efficiency cannot offset rising pricesinsightam a big fanboy for Astra/Codex right now but anyone claiming that token efficiency makes up for continuously DOUBLING PRICES is either delusional or lying to you this has been a trend since GPT 5: - GPT‑5: $1.25 in / $10 out - GPT‑5.2: +40% in / +40% out - GPT‑5.4: +43% in / +7% out - GPT‑5.5: +100% in / +100% out - GPT‑6 Astra: +100% in / +67% out Astra is now $10 in / $50 out that’s 8x on input and 5x on output no token efficiency in the world makes up for this and let’s not start talking about the priority multiplier that quietly went from 1.5x to 2x to a whopping 2.5x and recently back down to 2x OpenAI is in no way shape or form “cheap” not saying this to dunk on them quite the opposite this only shows that they’re confident that they’ve positioned themselves where they don’t need to discount anymoreRobin EbersSep 2627Fable 5’s re-release changes model accesstipON FABLE 5 RE-RELEASE Thariq clarified that 1) the new classifiers shouldn't happen as often as the blog post might have suggested 2) it will be fully transparent when your request was rerouted (in logs I guess?) 3) you will NOT be charged Fable rates when a request was rerouted to Opus this is a 10/10 update and how I wish Anthropic was always communicating GG!Robin EbersJul 2619AI changes what software teams can buildinsighti know this sounds crazy because GPT 5.6 Sol is a great model but if i had to choose between 5.6 Sol and Opus 4.8 i’d still pick Opus hear me out on this i promise this is not bait claude models remain the best at writing, design, and “jack-of-all-trades” intelligence. they’re still the best at everyday tasks and obviously OpenAI knows this GPT 6 is now rumored to be rushed to the finish line to better compete with Anthropic for probably that exact reason if all models would disappear overnight and all i had was Sol, i’d be extremely happy. but the model isn’t quite the jump that Fable was, and it still lacks behind across the board explicit examples include model behaviors too GPT 5.x series models have this weird habit of digging themselves into a hole that is very hard to convince them to get out of the only model doing this even more is Grok Opus and Fable feel more natural, more agreeable, more eager to understand YOU and think about different angles and despite having a $200 Codex sub, i still spend my time running down Fable in Claude, and switching to Opus for less important tasks instead of using GPT 5.6 it’s been a few days now and i’ve tested them A LOT i’ll obviously continue to test and use them, and i’m impressed what OpenAI was able to squeeze out of the old model base but it’s time to move onRobin EbersJul 2618Cursor Sand can triage Slack, Sentry, and Linearguideso this seems to be what was previously rumored to be "Cursor Sand" testing this, still early, but one thing I just set up: 7am every morning - check slack channel for client feedback - check sentry for new issues - check linear for issues - triage between all three - validate in code - propose top fixes - delegate to cursor cloud pretty wildRobin EbersAug 2613Grok 4.6 and Fable make a strong pairinsightGrok 4.6 + Fable 5 is all you need these two together are a killer combo 4.6 is especially great because it is less eager to finish the task. it can actually run for a while and check its own work. if 4.7 truly is that much better (as Elon claims), i’d say OpenAI is in big trouble Anthropic is fine, their ecosystem is much stronger and can absorb more pricing and model volatility really curious about Fable 5.1 thoughRobin EbersAug 268Codex can work autonomously for long sessionstipCodex is really doing this? 3 days into letting Codex build a local World of Warcraft clone it's pretty insane so far frame rate is poor, but it's been working for a few hours to fix that 🤣 https://t.co/wUk3FSHZUsRobin EbersJul 265
Read post