August 12, 2026
Korean AI lab Upstage has released Solar Pro 4, scoring 42 on the Artificial Analysis Intelligence Index, a significant increase from Solar Pro 3’s 14
See model pageSolar Pro 4 is Upstage AI's new proprietary flagship reasoning model, replacing Solar Pro 3 from April 2026. At 42 on the Intelligence Index it sits alongside Inkling (xhigh, 42) and just behind MiMo-V2.5-Pro (43), and shows a 27-point increase over Solar Pro 3. Pricing increases to $0.30/$1.20/$0.06 per 1M input/output/cache hit tokens from Solar Pro 3's $0.15/$0.60/$0.02 via Upstage’s first-party API.
Key results:
➤ Solar Pro 4’s largest improvements on Solar Pro 3 are on agentic and long context work. Terminal-Bench v2.1 improves from 12% to 57%, AA-LCR from 31% to 71%, and τ³-Banking from 9% to 23%. GDPval-AA v2 shows strong progress on real-world agentic tasks, where Solar Pro 3 scored an Elo of 498, well below the human baseline of 1000, Solar Pro 4 scores 1277.
➤ AA-Omniscience improvement from -53 to -1 was from abstaining on more questions. Solar Pro 4 attempts only 41% of questions against 92% for Solar Pro 3, and its hallucination rate is 24%, higher than Command A+ (14%) and MiniMax-M3 (18%), and a vast improvement from Solar Pro 3’s 88%. AA-Omniscience Accuracy remains unchanged at 19%.
➤ Solar Pro 4 is more token efficient than Solar Pro 3, though still verbose for its intelligence level. It uses 43k output tokens per Intelligence Index task, around 17% fewer than Solar Pro 3's 52k.
➤ The intelligence gain comes with a hit to latency. Solar Pro 4 takes 8.6 minutes to complete an average Intelligence Index task, against 6.0 minutes for Solar Pro 3, despite using fewer output tokens per task.
➤ Pricing is $0.30/$1.20 per 1M input/output tokens. This is in line with MiniMax's first-party pricing for MiniMax-M3, which scores 3 points higher at 45, and is more expensive than DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at $0.14/$0.28 and 52 on the Intelligence Index. Cache hits are priced at $0.06 per 1M, an 80% discount on input token price.
Additional model details:
➤ Context window: 384K tokens
➤ Max output tokens: 256K
➤ Modalities: Text input and output only
➤ Pricing: $0.30 / $1.20 / $0.06 per 1M input/output/cache hit tokens
➤ Inference providers at time of launch: Upstage first-party API, OpenRouter

Solar Pro 4's strongest improvement is on real-world agentic work. It scores an Elo of 1277 on GDPval-AA v2, up from 498 for Solar Pro 3. Solar Pro 4 sits above the human Elo baseline of 1000, slightly ahead of Qwen3.7 Max (1272) and MiMo-V2.5-Pro (1266).

Solar Pro 4's AA-Omniscience score improves from -53 to -1, however the improvement comes from abstention rather than knowledge. It attempted only 41% of questions compared to 92% for Solar Pro 3, and while its hallucination rate improves from 88% to 24%, its AA-Omniscience Accuracy is almost unchanged at 19%.

Solar Pro 4 uses 43k output tokens per Artificial Analysis Intelligence Index task, ~17% fewer than Solar Pro 3's 52k. However, it is still verbose compared to other models for its intelligence level.

Solar Pro 4 takes 8.6 minutes per Artificial Analysis Intelligence Index task, up from 6.0 minutes for Solar Pro 3, despite using fewer output tokens per task.

Full results across across the Artificial Analysis Intelligence Index:

Read the latest

Intelligence at pocket scale: Benchmarking small models and mobile phones
Independent intelligence benchmarking of small language models on a set of evaluations chosen for mobile device use, launched alongside mobile phone inference benchmarking with Liquid AI. We evaluate the same quantized builds used on mobile phones, and performance is measured on real devices.
August 24, 2026

Announcing the Speech Agent Arena: Compare Speech agents in real world conversations
Announcing our new Speech Agent Arena, evaluating Speech to Speech models on real-world scenarios to analyze conversational preference and task success rate
August 24, 2026

Announcing the Artificial Analysis Search Index: Same Agent, Different Search
The Artificial Analysis Search Index benchmarks how search API providers perform on quality, cost, and speed when used by an agent. We compare different Search API providers across a series of search-related benchmarks using the same agentic setup.
August 18, 2026