Note If you’re here for the recipe and don’t want the exposition: Agent Instructions “A computer can never be held accountable; therefore a computer must never make a management decision.” - IBM training manual, 1979 Until recently, the primary concern of delegating judgment to the machine has been outward-facing, directed at the individuals and societies left to the whims…
I’ve been working on a large refactor in one of my side projects - migrating from a tightly-coupled, hardcoded project barely more than a demo to a pluggable, adapter-based architecture. I’ve been using spec-driven development and helicopter-engineering the agents, so I’m fairly confident in the code. But there have been a lot of changes, so I want to make sure it works,…
AI tools change fast. This post is intended to snapshot how I use AI today and provide some context for how I arrived here. I do not anticipate keeping this post up-to-date, though I may revisit the idea with update posts in the future. Three years of history Tip Feel free to collapse this historic context if you just want to get to the tools and way of working! Note This narration is from memory,…
LinkedIn is rife with the drive-by copy-paste of raw LLM output, and the trend has bled into Discord, Reddit, and other online forums, work chats, and emails. I’m coining this ‘sloppypasta,’ and this is my rant against it. sloppypasta Verbatim LLM output copy-pasted at someone, unread, unrefined, and unrequested. From slop (low-quality AI-generated content) + copypasta (text…
“Treat them like new interns” is the common wisdom when leveraging LLMs to do any more-than-slightly-complex task. Chatting with LLMs is like writing on a whiteboard; conversations may fill the board, but they are wiped clean each time. Which begs the question – What if the LLM (or AI system built around it) could learn? Broadly, there are two approaches to this. The first…
In a recent episode of the Latent Space podcast ( Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith ), the Artificial Analysis team pointed out a strong correlation between model performance on their AA-Omniscience Accuracy benchmark and the model’s parameter count: An interesting thing about this accuracy metric is that it tracks more closely…
These are my predictions for AI (primarily LLM-focused) in 2026 along with my certainty / probability estimates. Agents, Cont’d Agents aren’t going anywhere. Their capability will increasingly be driven by Reinforcement Learning (RL). We will see a divergence: domains with verifiable tasks (math, coding) will advance rapidly because the reward signals are clear. Subjective domains will…
Back in January, I made a series of predictions for 2025 , assigning somewhat arbitrary probability estimates to each as indicators of my confidence in the prediction. Now that the year is wrapping up, it’s time to see what kind of AI Nostradamus I am. Agents, agents, agents Prediction: We will see an even stronger push for Agents in 2025. Probability: 100% Verdict: Correct ✅ This was a…
Until recently, I held the opinion that training custom language models was inadvisable except in relatively rare cases. Simply requiring marginally better performance on your task was insufficient justification for customization; subsequent generations of models from foundation labs would inevitably catch up. Only when you had extreme latency constraints, specific tasks with low drift, and/or…
I’ve recently been enamored with the concept of “fractal complexity”, not pertaining to the mathematical definition of “complexity”, but rather as a way to explain why “it depends” is such a frequent response in conversations between people with various levels of expertise. I believe I came across this idea while listening to an episode of “Brad and…
I had intended to start this post by proclaiming “2026 will be the year of reinforcement learning” as 2025 is “the year of agents”… But model and research releases over the past several weeks indicate that it might be that “H2 2025” is when reinforcement learning for agentic AI really takes off. Open model releases such as Qwen3 (particularly the recent…
Sycophancy On April 25th, we [OpenAI] rolled out an update to GPT‑4o in ChatGPT that made the model noticeably more sycophantic. It aimed to please the user, not just as flattery, but also as validating doubts, fueling anger, urging impulsive actions, or reinforcing negative emotions in ways that were not intended. 1 Sycophancy ( click here for pronunciation guide/recording ) is overwhelming,…
On a recent Dwarkesh podcast, Sholto Douglas, an AI researcher working on reinforcement learning at Anthropic said: “There is this whole spectrum of crazy futures. But the one that I feel we’re almost guaranteed to get—this is a strong statement to make—is one where at the very least, you get a drop-in white-collar worker at some point in the next five years. I think it’s very…
I’ve been saving up for a year and change for a GPU upgrade, and my local Microcenter finally had 5090s in stock. So I bought one. It is a CHONK – I’m getting full “you vs. the guy she tells you not to worry about” vibes: Top view comparing 3090 and 5090 (5090 is larger) Side view comparing 3090 and 5090 (5090 is larger) Specifically, I’ve upgraded from a PNY…
I previously posted an introduction to the AI Treadmill (so named because trying to keep up with everything felt like running on a treadmill cranked all the way up) and provided a data dump of the list of links I shared over the course of 2024. Earlier this month, I made those (and all Treadmill updates since) available on this site. Look for the AI Treadmill on the homepage or in the top right…
Superintelligence Strategy is a policy paper by Dan Hendrycks (Director, Center for AI Safety), Eric Schmidt (former CEO, Google), and Alexandr Wang (CEO, Scale AI), that proposes a three-part framework to manage the risk associated with Artificial SuperIntelligence (ASI). Strategies include Mutual Assured AI Malfunction ( MAIM – shouldn’t “Mutual Assured AI Malfunction” be…
The Bitter Lesson proposes that general-purpose solutions based on the pace of technological progress and scale ineveitably outperform previously state-of-the-art but handcrafted solutions. In many cases, the Bitter Lesson might be considered optimistic - technological progress drives a positive feedback loop of improvement. What is left unsaid in that optimism is the pessimistic implication…
Facebook’s fasttext and Google’s cl3d ) were both deprecated mid-2024. While multilingual LLMs can accomplish language identification tasks, using even a 7B-parameter LLM to determine a text’s language is overkill. I set out to find a replacement that is at least as fast and performant as the deprecated models. Code from these experiments is available here . TLDR Lingua is the…
I’ve recently started using uv to manage my python dependencies. Previously, I primarily managed my dependencies with conda , and supplemented with pip when packages were not available through anaconda.org. This process update is the result of workflow friction when using conda - it was unpleasant to jump through the required hoops myself, and unreasonable to expect others to do it when had…
Here are my predictions for 2025 along with my certainty / probability estimate: Agents, agents, agents We got nascent function-calling and JSON mode toward the end of 2023, which morphed into formal structured generation over the course of 2024. 1 2 Combined with an explosion of Agentic frameworks ( LangChain , LangGraph , Semantic Kernel , AutoGen / AG2 , CrewAI , PydanticAI , etc.), 2024 was…
At the end of 2023, I was transferred to the team working on PMI Infinity , a (then-nascent) AI-powered tool and product for Project Managers, to act as an AI/ML Engineer and provide expertise on the AI/ML aspects of developing a product based on generative models. Coming from more standard data science and machine learning, my impostor syndrome flared and I felt behind the 8-ball. I’d been…
While I generally attempt to avoid pendantry, I am acutely aware that word choice – and the specific connotations and denotations of the selected terms – matter. Therefore, I strive to be specific in the language I use, and it niggles my noodle a bit when words or phrases are used incorrectly. However, as a Libra, I am apparently celestially obligated to desire balance and equilibrium,…
This is part three of a three-part series ( one , two ) where I explore best practices for evaluating RAG architecture via Ragas’ recent v0.2 release (specifically, Ragas v0.2.3 ). Code from these experiments is available here . In this post, I’ll use Ragas to investigate my hypothesis - LLMs prefer answers generated by themselves (as opposed to being objective evaluators). This is,…
This is part two of a three-part series ( one , three ) where I explore best practices for evaluating RAG architecture via Ragas’ recent v0.2 release. Code from these experiments is available here . In this post, I will dive into why I’m so excited about Ragas v0.2 and dive into how it works. Specifically, the I am referencing Ragas v0.2.3 ; the team is rapidly iterating and this…
This is part one of a three-part series ( two , three ) where I explore best practices for evaluating RAG architecture via Ragas’ recent v0.2 release. Code from these experiments is available here . This post covers the preliminary / background material. In later posts, I’ll cover what makes Ragas v0.2 so special, how it works, and run an experiment with it. What is RAG? Retrieval…
A few weeks ago at work, I wanted to ensure that the prompt template we used with Semantic Kernel transformed into the OpenAI API spec messages array that I expected. Little did I know that this simple objective would take me a few days, several experimental notebooks, a thorough tour through the Semantic Kernel’s dev blog and (limited) documentation, and a review of pretty much the entire…
Over the past month, I’ve seen several articles reporting on the dangers of “model collapse” popping in news aggregators (HackerNews, Reddit, etc.). 1 2 3 As someone working in the space, I find the timing of these reports interesting because research papers naming the phenomenon (“model collapse” 4 or “model autophogy disorder” 5 ) came out over a year…
The primary training task for generative AI models is inherently plagiaristic. New AI-based answer engines leverage these plagiaristic models to provide in-engine responses containing the specific information or content the user is looking for, obviating the need for the user to click through to the (original) content and depriving the creator of traffic. This breaks the content…
This is part four of a four-part series ( one , two , three ) where I examine the influence typos have on LLM response quality. Code from these experiments is available here . In this post, I induce typos in a standardized set of prompts with increasing frequency in the hopes of understanding how typos influence generation . If you recall, my hypothesis is that typos will increase error rates…
This is part three of a four-part series ( one , two , four ) where I examine the influence typos have on LLM response quality. Code from these experiments is available here . In this post, I induce typos in a standardized set of prompts with increasing frequency in the hopes of understanding how typos influence meaning as represented by sentence embeddings . Note An embedding is the vector (i.e.,…
This is part two of a four-part series ( one , three , four ) where I examine the influence typos have on LLM response quality. Code from these experiments is available here . In this post, I use the typo generation function to induce typos with increasing frequency in the hopes of understanding how typos influence tokenization . Recall my hypothesis: Typos increase token counts – as the…
A recent theme in conversations at work is that “prompts are fragile.” Word choice and word order can have large impacts on LLM responses, and every user input is a potential attack vector for the LLM equivalent of a SQL injection attack. But before I go off on that tangent, having this repeated discussion got me thinking – “Can I quantify how sensitive LLMs are to…
Meta provided insight into some of the costs of training LLMs in their Llama 2 1 and Llama 3 2 papers, listing GPU hours and GPU power draw required to train the models, and t C O 2 CO_{2} C O 2 eq , or metric tons of Carbon Dioxide Equivalent emitted. Of note, we see a huge generational increase from Llama2 to Llama3. It took 12x the electricity to train Llama3-8B as it did to train Llama2-7B,…
The “open weights” or “open model” LLM ecosystem is thriving, with major new releases from Meta, Microsoft, Databricks, and Snowflake in the past two months. Given that Meta’s Llama 2 family became a standard for comparison for all other models, I thought it would be useful that aggregate all of the information of the ‘Llama 3 cohort’ in a single place.…
Web crawlers & search How do sites show up in search results? Search engines (among other actors) run web crawlers – bots that explore websites so the search engines (among others) to catalog and index site contents. This index helps search engines to provide relevant content when you search. In the ’90s, webmasters noticed that crawlers could place heavy loads on web servers. They…
Yann LeCun believes autoregressive models inherently suffer from a kind of ‘generational drift’ 1 due to compounding errors. Coming from LeCun - Chief AI Scientist at Meta and famed for his work on convolutional neural networks and image recognition - this alarming statement carries a lot of weight. It is also, I think, a novel concept to people who are not steeped in statistics, or…
I’m Alex Graber, an AI/ML Engineer at the Project Management Institute where I work on Infinity , an AI-powered tool and product for Project Managers. Prior work includes Data Science and ML positions where I focused on business analytics and MLOps. In which I organize my thoughts I often email myself about an incipient notion hoping that I’ll be able to revisit it later from a harried…