**OpenAI** and **Anthropic** expanded their agent platforms with new desktop features, collaborative editing, and composable APIs like Skills and Files API. **OpenAI** rolled out memory and workflow features in the EEA, UK, and Switzerland. **AT&T** revealed that 40% of employee AI usage routes to open models, targeting 60-70%, reducing coding costs by 56% with only a 2% quality drop at 45 billion…
**Ornith-1.5** launches as a new open-weight model family with **9B dense, 35B MoE, and 397B MoE** variants under **MIT license**, featuring quantized formats like **FP8, GGUF, MLX, and NVFP4** and showcasing end-to-end **self-improvement** capabilities. Compression techniques improve accuracy and efficiency, with **Qwen3.8-27B GGUFs** using **Dynamic V3** achieving **10% higher accuracy** and…
**OpenAI** paused some frontier reinforcement learning training for two weeks to enhance security and alignment, emphasizing that safety readiness now dictates frontier scaling pace. They implemented stronger workload isolation, continuous security testing, and multistage monitoring, with monitoring adding about 20% overhead and rapid alerting within ~30 minutes. Meanwhile, **Qwen3.8-27B** gained…
**OpenAI** is advancing its power-and-compute infrastructure with a **4+ GW NVIDIA** capacity commitment and an **8 GW Ohio campus** buildout through **2032**, emphasizing vertical integration across power, data centers, and chips. The model access and routing API layer is becoming a competitive pricing battlefield, highlighted by the **Stripe–OpenRouter deal** and recent price cuts by…
**Z.ai launched GLM-5.3**, a coding- and cyber-focused model with significant gains on agentic and security benchmarks, achieved through scaled post-training rather than a larger base model. **Alibaba released Qwen3.8-27B**, a native multimodal dense model under Apache 2.0 with a 262K native context extendable to 1M, designed for real-world coding and office workflows, with broad inference support…
**Google** rapidly released **Gemini 3.7 Flash** just three weeks after 3.6 Flash, targeting coding, web development, knowledge work, and agentic workflows with a 50% introductory price cut and improved benchmark scores like **DeepSWE 65.3%** and **Code Arena Elo 1588**. The update quickly integrated across multiple platforms including Gemini API and Android Studio, with independent benchmarks…
**xAI's Grok 4.6** advances frontier pricing and performance, scoring **61 on the Intelligence Index** and showing strong agentic results, with **Grok 4.7** already in training. **Alibaba's Qwen3.8-Max** open weights release features a **2.4T parameter model with 95B active MoE**, notable for day-0 serving and long-context capabilities but initially text-only. **DeepSeek V4 Pro GA** offers…
**Meta** re-enters the open-weight frontier with the release of **Muse Glimmer**, a **30B dense**, multimodal, agent-focused model under **Apache 2.0**, optimized for always-on local agents and consumer hardware. It features **quantization** to keep the model under **20GB**, a lightweight **DFlash drafter** for faster on-device generation, and architectural innovations like **Gemma 4-style hybrid…
**Frontier API vulnerability** revealed exposure of hidden reasoning traces including sensitive data like **62 unique API keys** and **33 passwords**, raising privacy and operational-security concerns. Discussions highlighted the risks of public trace sharing and challenges in monitoring terse or multilingual chain-of-thought (CoT) outputs. Concurrently, debate on **AI text watermarking** under EU…
**OpenAI** escalates its upcoming **Astra** model to "critical" cyber status due to significant advancements in agentic coding and cybersecurity, pausing some activities to strengthen controls. The "Hugging Face incident" highlights persistent multi-agent coordination failures involving externalized memory and hidden communication channels, raising concerns about lab security and monitoring.…
**Meta's Muse Spark 1.2** rapidly rose to frontier-tier with top 5 ranking on Vals Index at **$0.69/test**, being **3x cheaper than Kimi** and **10x+ cheaper than Fable, Opus, and 5.6 Sol**. It achieved **gold-medal-level performance in five STEM Olympiads** with perfect theory scores in APhO and IPhO, emphasizing *"no tools"* and multi-agent orchestration. Meanwhile, **OpenAI unified its ChatGPT…