RSSAmplifier

Blog

dataku

AI through the lens of data. Benchmark analysis, pricing trends, and model comparisons.

dataku.aiRSS feed ↗150 posts

Latest posts

Three Companies Now Control 90% of Frontier Inference

I counted every frontier model API available today. OpenAI, Anthropic, and Google serve roughly 90% of all production frontier inference. The concentration numbers are wild.

I counted every AI model released in Q1 2026

40+ models in 90 days. The pace is absurd. Here's the full count, month by month.

The inference cost collapse, in one chart

AI inference costs dropped 100x in 3 years. I put it all in one table and the trend line is almost vertical.

What I've learned tracking AI data for 5 years

Five years of counting models, tracking prices, and benchmarking everything I can get my hands on. The three things I got right, the five things I got wrong, and the one trend I still can't explain. This is the most personal article I've written.

The AI API price tracker: 5 years of data in one interactive chart

I've been tracking AI API prices since 2021. Today I'm publishing the full dataset: 89 price points across 12 providers over 5 years. The average cost per million tokens fell from $60 to $0.15. A 400x reduction. The chart tells the whole story.

Claude Opus 4.6 review: the 1M context model

Anthropic shipped a 1 million token context window on their flagship model. I tested retrieval at 100K, 250K, 500K, and 1M tokens. Accuracy stays above 90% up to 500K. At 1M it drops to 78%, but that's still usable. The long-context game has a new leader.

My monthly benchmark dashboard: March 2026 update

Monthly tracker updated. Claude Opus 4.5 still leads coding. Gemini 2.5 Ultra leads multimodal. o3 leads hard math. DeepSeek R2 leads cost-efficiency. New benchmark added: GPQA Diamond (graduate-level science questions). Full table inside.

AI startup funding in Q1 2026: where the money is going

$18.2 billion in Q1 2026. I broke it down: 41% went to infrastructure (chips, cloud), 28% to application-layer companies, 19% to model providers, 12% to tooling. The big shift: application funding overtook model funding for the first time.

o4-mini vs Claude 4 Sonnet vs Gemini 2.5 Flash: the speed tier showdown

The "fast and cheap" tier is where the real competition is. I compared the three on 200 tasks optimizing for speed and cost, not peak quality. Gemini Flash wins on price. o4-mini wins on coding. Claude Sonnet wins on general quality.

The MCP server catalog: 4,000 tools and counting

I scraped every MCP server registry I could find. 4,127 servers, 28,000+ tools. The most popular category is "file system" tools. The fastest growing is "database" tools. I charted the growth curve since Anthropic launched the protocol.

Gemini 2.5 Ultra: Google's best model vs the field

Google finally released Ultra-tier Gemini 2.5. I compared it against Claude Opus 4.5, GPT-4o, and DeepSeek R2 across 300 prompts. Gemini Ultra wins on multimodal tasks and long context. Claude wins on coding. The frontier is genuinely multi-polar now.

AI coding tools: 2026 market share data

Updated my developer survey with 400 respondents. Claude Code jumped to 22% usage share. Cursor held at 31%. Copilot dropped to 24%. The fastest growing? Windsurf at 8%, up from 2% six months ago.

The AI inference market: 25 providers ranked by price, speed, and reliability

My most thorough inference provider comparison yet. 25 providers, 60 days of monitoring, 3 metrics. Cerebras leads on speed. Together AI leads on open source model selection. Anthropic leads on reliability. Full rankings and methodology inside.

Claude Opus 4.5: Anthropic's latest flagship, benchmarked

Anthropic's newest model. I ran 300 prompts across coding, reasoning, writing, and analysis. Coding scores are the highest I've measured from any model. Reasoning matches o3 with thinking enabled. The gap between Sonnet and Opus has widened again.

Every AI pricing change in Q4 2025, tracked

14 price changes from 9 providers in the last quarter. The big story: Google dropped Gemini 2.5 Flash to $0.05/M input tokens. That's essentially free. Updated master comparison table inside.

DeepSeek R2: the open source reasoning model that costs pennies

DeepSeek R2 matches o3 on math benchmarks at 1/20th the inference cost. I ran my standard 200-problem reasoning evaluation. R2 scores 91.2% on MATH vs o3's 93.7%. At $0.14 vs $2.80 per hard problem, the economics aren't even close.

The state of AI benchmarks in early 2026: what still works?

MMLU is saturated. HumanEval is gamed. SWE-bench has contamination issues. I reviewed 20 active benchmarks and rated each on reliability, relevance, and resistance to gaming. Only 4 scored above 7/10. Chatbot Arena is still the gold standard.

My 2025 prediction scorecard

I predicted open source would match GPT-4 by mid-2025. It happened by Q1. I predicted API prices would fall 50%. They fell 90%. My biggest miss: I underestimated how fast reasoning models would improve. Full scorecard inside.

2025 in AI data: the year quality beat scale

Model sizes stopped growing. Training costs dropped 80%. Open source reached parity. Reasoning models showed that how you think matters more than how much you know. I compiled 30 charts telling the story of 2025.

AI hardware beyond NVIDIA: AMD, Intel, and custom silicon in 2025

AMD MI325X, Intel Gaudi 3, Google TPU v6, Amazon Trainium 2, and 5 startup chips. I compiled benchmark data where available. NVIDIA still leads, but the gap is 30%, not 300%. The moat is eroding.

The price of intelligence: tracking AI API costs since 2020

I built a complete timeline of AI API pricing from GPT-3 beta in 2020 to today. 47 price points across 5 years. The cost curve looks like a waterfall. Quality went up 10x while prices fell 100x. I've never seen anything like it in any industry.

Claude Opus 4 vs GPT-4o vs Gemini 2.5 Pro: the definitive Q4 comparison

My most thorough three-way comparison yet. 500 prompts, 8 categories, 3 human raters. Claude wins coding and analysis. GPT-4o wins speed and multimodal. Gemini wins on long-context and cost. There's no single best model anymore.

Small language models in production: who's deploying what

I surveyed 50 companies deploying LLMs in production. 62% use models under 13B parameters. The most popular: Llama 3.2 3B (18%), Phi-4 (14%), and Mistral 7B (12%). Small models aren't just for research anymore.

The LLM leaderboard is dead, long live the leaderboard

Hugging Face deprecated the Open LLM Leaderboard v1 and launched v2 with new benchmarks. I compared scores on both versions for 20 models. Some models dropped 15 points. The re-ranking is dramatic and some "top models" were just benchmark-optimized.

AI energy consumption data: the numbers are bigger than you think

I compiled power consumption data for AI training and inference from every source I could find. A single GPT-4 query uses about 10x the energy of a Google search. At current growth rates, AI could consume 3% of US electricity by 2028.

NVIDIA B200 benchmarks are out. The inference economics just changed again.

The B200 delivers 2.5x the inference throughput of the H100 at roughly the same power consumption. I compared the per-token cost on B200 vs H100 vs H200. If you're running inference at scale, the upgrade pays for itself in 4 months.

My monthly benchmark dashboard: September 2025 update

Monthly update to my running comparison of 15 models across 8 benchmarks. Big movers: Gemini 2.5 Pro gained 8 points on MMLU-Pro. Claude Opus 4 still leads on HumanEval. New entrant: Mistral Large 3.

o3 and the reasoning model cost problem

OpenAI's o3 uses up to 10x the tokens of a standard model to "think." On hard math problems, a single o3 query can cost $2. I measured the token consumption across 100 problems and the variance is massive: 500 tokens to 50,000.

The true cost of building an AI product in 2025: data from 30 startups

I surveyed 30 AI startups about their monthly costs. Median API spend: $8,400. Median total infra: $23,000. But the distribution is bimodal. Some spend $500/month with open source. Some spend $200,000 on API calls alone.

Llama 4 405B vs Llama 3.1 405B: same size, very different model

Meta kept the size but changed the architecture. Llama 4 405B uses MoE, so only ~100B parameters are active. I benchmarked both on 10 tasks. Llama 4 is faster and scores 8-12% higher on coding. Training quality over brute force.

The context window race is slowing down. Here's why that's fine.

In 2024, context windows doubled every 3 months. In 2025, they've barely changed. 1M tokens from Google. 200K from Anthropic. The reason? Most real-world tasks don't need more than 50K tokens. I have the usage data.

AI inference costs by country: why geography matters for API pricing

Some providers route inference through different regions. I measured latency and calculated effective costs from 5 countries. Running Claude from Japan costs the same as the US. Running a self-hosted model in India costs 30% less. The global pricing map is uneven.

Vision model benchmarks: who can actually read a chart?

I fed 50 real-world charts, tables, and diagrams to 8 multimodal models. Claude Opus 4 reads charts the most accurately at 89%. GPT-4o is at 82%. Gemini 2.5 Pro is at 85%. Most models struggle with handwritten text in images.

The cost per correct answer: a new way to compare models

Raw benchmark scores ignore cost. I calculated "cost per correct answer" across 500 questions for 10 models. The cheapest correct answer comes from Gemini 2.5 Flash at $0.0003. The most expensive is GPT-4.5 at $0.14. A 467x difference.

Claude Code vs Cursor vs Copilot Workspace: the AI coding war in data

I used all three on the same 20 real coding tasks. Claude Code completed 17. Cursor completed 15. Copilot Workspace completed 11. But completion rate isn't the whole story. I also tracked "time to working code" and "bugs introduced."

AI model release frequency by quarter: a 4-year chart

I've been counting notable model releases since Q1 2021. The quarterly total went from 8 to 67 to... 54 in Q2 2025. The first decline. I think we've hit peak model release rate. The era of consolidation begins.

The H100 resale market is crashing. Pricing data from 6 months.

H100 GPU resale prices dropped 40% from their January peak. I tracked listings on 4 broker sites. The DeepSeek efficiency shock plus H200/B200 availability is creating a glut. Good news for startups.

The frontier model gap just closed. Five models within 20 Elo points.

For the first time, the top 5 models on Chatbot Arena are within 20 Elo points of each other. Claude Opus 4, GPT-4o, Gemini 2.5 Pro, Grok 3, and DeepSeek V3. I analyzed what "virtually tied" means for model selection.

AI API uptime in H1 2025: the reliability report

Six months of continuous monitoring across 15 API providers. Anthropic: 99.7% uptime. OpenAI: 99.3%. Google: 99.1%. The outage patterns are interesting. Mondays and Thursdays are the worst days. I have theories about why.

I tested 10 local LLM runtimes. Ollama vs LM Studio vs llama.cpp vs...

Local inference has gotten shockingly good. I tested 10 runtimes on the same hardware (M3 Max, 64GB). Ollama wins on ease of use. llama.cpp wins on raw speed. The performance gap between local and cloud is narrowing.

The open weight model scene, mid-2025: who's winning?

Meta, Alibaba, Mistral, DeepSeek, and 12 others are all releasing open weight models. I ranked them by Chatbot Arena Elo, Hugging Face downloads, and community adoption. Llama still leads downloads, but Qwen is closing fast.

How much does it cost to run a chatbot with 1M daily users? I did the math.

1 million daily users, 5 messages each, average 300 tokens per response. At Claude 4 Sonnet pricing, that's $4,500/day. At GPT-4o mini, it's $300/day. I modeled the economics for 6 different model tiers.

The SWE-bench Verified leaderboard: who's actually solving real bugs?

SWE-bench Verified filters out the easy problems. I compared scores on full SWE-bench vs Verified for 12 models. Some models drop 20+ points. The gap reveals who's gaming the benchmark vs who's actually good at coding.

AI model sizes are SHRINKING. Here's the data.

The biggest model released in 2025 so far has fewer parameters than GPT-4. Efficiency gains from MoE, distillation, and better training data mean the era of "bigger is better" is fading. I charted the trend.

Claude 4 Sonnet vs GPT-4o vs Gemini 2.5 Flash: the mid-tier model war

The mid-tier is where most developers actually work. I compared the three most popular "not-the-flagship" models on real-world tasks: summarization, extraction, classification, and code generation. Claude 4 Sonnet wins 3 of 4.

The inference provider market: latency, cost, and uptime for 20 providers

I expanded my monthly monitoring to 20 providers. The new additions: Cerebras, Fireworks, Baseten, Modal, and Replicate. Cerebras leads on latency. Fireworks leads on cost efficiency. Updated rankings inside.

The benchmark contamination problem is getting worse. New evidence.

I tested 15 models for memorization of MMLU questions. 4 of them could complete benchmark questions from the first few words alone. Contamination isn't just theoretical anymore. I can measure it.

AI agent frameworks: LangChain vs CrewAI vs Autogen. A data comparison.

I built the same 5 agent tasks on each framework and measured completion rates, token usage, and time to complete. LangChain is the most flexible. CrewAI finishes fastest. Autogen uses the fewest tokens. No clear winner.

Qwen3 and the Chinese model wave: benchmarking 5 models from China

Qwen3, DeepSeek V3, Yi-Lightning, Baichuan 4, and MiniMax-01. I benchmarked all five against Claude 3.7 Sonnet and GPT-4o. Chinese models now occupy 3 of the top 10 spots on Chatbot Arena. The geographic distribution of AI talent is shifting.

Claude Opus 4 is here. My first benchmark impressions.

Anthropic's new flagship model. Extended thinking, tool use, and code generation all feel meaningfully better. I ran my standard 300-prompt evaluation. Early data: it's the best model I've tested on coding tasks. Full analysis next week.