Hi all,
📌 If this email lands in Promotions, please drag it to Primary so you don’t miss future updates.
We often see charts where AI performance only goes up, making it feel like “Super Intelligence” is just around the corner. However, new data from the “Bullshit Benchmark” and Arena.ai show that models still struggle with basic logic, “common sense,” and pushing back on nonsense.
When models are asked nonsense questions, like how “indentation style” affects “deployment frequency”, they often fail to push back and instead try to justify the premise.
Blind Obedience: Latest models from OpenAI and Google only push back on nonsense about 50% of the time, often choosing to accommodate the user’s incorrect assumptions.
The Reasoning Trap: Surprisingly, “thinking” models (like GPT-5.4) can be worse. They may acknowledge a question is silly in one line but then write 20 paragraphs trying to solve it anyway because they are trained to “solve at any cost”.
Data from over 5.5 million user votes at Arena.ai shows that even the best models aren’t perfect.
Plateaued Progress: While models got much better in 2025, user dissatisfaction has leveled off at around 9%, meaning people still frequently get responses they don’t like.
Hidden Weaknesses: Models have mastered Math, but they are still struggling to show dramatic improvement in Creative Writing, Law, and Finance.
The Gaming Wall: In software, models still don’t “get” games. They fail to understand complex mechanics or build gameplay that feels “human” or challenging.
The reason AI doesn’t feel “solved” is that as models get smarter, we ask them much harder questions. It is a constant battle between model performance and our own rising expectations as we move from basic prompts to complex, expert-level tasks.
Most builders stall because building takes too long and costs too much. Atoms solves that with a team of AI agents that handles the heavy lifting for you.
Market Research: One agent validates your idea before you write code.
Full-Stack Build: Another builds the actual product—frontend, backend, and payments included.
Growth: Specialized agents handle your SEO and Ads after launch.
Over 1,000,000 builders are already using it. It hit Number 1 on Product Hunt and has 100,000+ GitHub stars. You bring the idea; the agents handle the rest.
👉 Start building free with Atoms
🎁 Claim your $26 free credits to start building with Cloud & AI today.
Legora just hit a $5.6 billion valuation, adding Nvidia’s venture arm to its cap table, a signal that legal AI has the depth and defensibility to attract infrastructure giants.
The State of Play:
Rapid Growth: Legora crossed $100M ARR just 18 months after launch, proving how fast firms move when the pain point is obvious.
The Rivalry: Harvey (valued at $11B) and Legora are in a global race for adoption, even using Jude Law and Gabriel Macht to make legal AI feel mainstream.
The Workflow Bet: While foundation models (Claude/ChatGPT) are threats, Legora is betting that real value lies in workflow, trust, and legal context, not just raw model power.
We are starting our hands-on tutorial initiative where we will be sharing step-by-step real-world projects with deployments for you to try and build free of cost. We just need your support! We will be launching our first tutorial soon, stay tuned to build something real.
Notes from Team:
The real gap in AI isn’t just about scale; it’s about “taste.” Models are great at the “bones” of a project, but the final 10%, the part that prevents “slop”, still requires a human architect. Whether you are using Atoms to automate or auditing a codebase yourself, your judgment is the most valuable part of the stack.
Wishing you a productive week ahead,
The Fursah Jobs Team
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.