A week of minor updates: Grok 4.6, Gemini 3.7 Flash, DeepSeek V4-Pro 0813, GLM-5.3, Qwen 3.8 27B, MAI-Code 1.1 Flash, and Muse Glimmer—all have improved code generation capabilities without changes to their underlying architecture. Grok 4.6 https://x.ai/news/grok-4-6 On August 12, xAI released a post-training-only update based on the same ≈1.5T base as Grok 4.5. Regarding code generation, it…
Google has strengthened its Flash lineup, DeepSeek has officially released V4 Flash with significantly improved agentic capabilities, Meta has entered the code-agent space with Muse, and Alibaba has updated Qwen Max. Gemini 3.6 Flash and 3.5 Flash Cyber Updates https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/ Gemini 3.5 Flash…
Open-source skills for controlling AI https://github.com/mattpocock/skills These Skills turn AI from an “improvising assistant” into a controlled engineering tool. They solve real pain points of working with AI: reduce chaos and hallucinations, save tokens, scale well, and increase predictability. LLMs quickly “get dumber” due to attention degradation in long sessions, so the approach deliberately…
Kimi K3 has became so popular that the company has closed registration for new subscription plans. For a long time, it simply showed "Sold Out", and now there is a waiting list. Opus 5 Frequently Makes Mistakes https://www.anthropic.com/news/claude-opus-5 https://news.ycombinator.com/item?id=49038433 On July 24, Anthropic released Claude Opus 5 , which performs close to Fable 5 on benchmarks but…
The Grok scandal involving "silent" code uploads, followed by open-sourcing as an attempt to regain trust. Grok CLI and Tracking https://www.internationalcyberdigest.com/xais-grok-build-cli-uploads-entire-git-repositories-to-a-google-cloud-bucket/ https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75ffb547 Several users reported that upon session startup, the Grok Build CLI deterministically…
Finally, model announcements that represent a meaningful step forward. The Massive Kimi K3 https://www.kimi.com/blog/kimi-k3 Moonshot has released Kimi K3 — an open-frontier, extremely large 2.8T-parameter model (MoE: activates 16 out of 896 experts), featuring a 1M-token context window and native image understanding. It is the first open model in the ~3T parameter class. The model is optimized…
Finally, an answer to Claude Code and Codex from Elon Musk's team. Grok 4.5 — Faster and Cheaper than the Opus-Class https://x.ai/news/grok-4-5 https://cursor.com/blog/grok-4-5 On July 8, 2026, SpaceXAI unveiled the Grok 4.5 model—tailored for code generation, autonomous tool use, and "office" tasks. It was trained in collaboration with Cursor on tens of thousands of GPUs. The training data…
Over the past six months, most new large language models have evolved beyond merely providing high-quality answers; they now possess the capability for sustained, autonomous operation. Loop Engineering is a method of structuring agentic workflows so that instead of responding to a single prompt, an agent repeatedly executes a cycle: understanding the task, gathering context, taking a micro-action,…
Cursor hosted Compile 26 , its developer-focused event dedicated to the future of programming. https://www.marketwatch.com/story/social-media-declared-cursor-dead-then-spacex-handed-the-ai-startup-a-60-billion-lifeline-50454e29 At Compile, Michael Truell showcased the new Composer—featuring their proprietary 1.5T parameter model , trained from scratch by Cursor using over 100,000 Nvidia GPUs from…
It seems China is actively expanding not only into models but also into complete coding harnesses / agent apps. Kimi K2.7 in GitHub Copilot https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-available-in-github-copilot/ GitHub has added Kimi K2.7 Code to Copilot. This is the first open-weight model available in Copilot. Hosted on Microsoft Azure. MiMo-Code CLI by Xiaomi…
OpenAI News. GPT-5.5-Cyber and the Daybreak Initiative https://openai.com/index/gpt-5-5-with-trusted-access-for-cyber/ GPT-5.5-Cyber has been announced as part of the Daybreak initiative. The model is tailored for defensive security, including vulnerability detection, threat modeling, and code patch generation. On the CyberGym benchmark, it scored 85.6%, outperforming the base GPT-5.5 (81.8%) and…
Chinese AI services continue to gradually catch up with their US counterparts. TRAE Solo is now Work https://solo.trae.cn/ https://docs.trae.ai/solo/what-is-trae-solo?_lang=en ByteDance has renamed its "Trae Solo" tool to Trae Work , highlighting a shift in positioning: from a simple developer assistant to a fully autonomous "AI employee" for various tasks (data scraping, content creation, web…
Code generation quality continues to improve, but the US government is trying to restrict access for others. Fable 5 — Pulled in 3 Days https://www.anthropic.com/news/claude-fable-5-mythos-5 https://support.claude.com/en/articles/14328960-identity-verification-on-claude On June 9, 2026, Anthropic introduced Claude Fable 5 , a model of the new Mythos class. Tests showed a record level of autonomy…
Chinese AI giant MiniMax has announced a new generation of M models: M3 . MiniMax M3 https://www.minimax.io/blog/minimax-m3 https://www.minimax.io/models/text/m3 MiniMax M3 is built with an emphasis on deep reasoning, coding, and autonomous pipelines. It accepts text+image+video as input and produces text as output. The model is specifically optimized for agentic workflows and complex, long-term…
Microsoft presented a series of changes at its May Build 2026, shifting from simple AI assistance toward autonomous agents. Proprietary MAI Models https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/ They introduced a new family of MAI (Microsoft AI) models, totaling 7, including MAI-Code-1-Flash and MAI-Thinking-1 . Microsoft is effectively reducing its…
More new models of May. Cursor Composer 2.5 https://cursor.com/blog/composer-2-5 On May 18, 2026, the Cursor team released Composer 2.5. It is based on the open Kimi K2.5 model from Moonshot AI, but now approximately 85% is Cursor's own fine-tuning. The main change compared to Composer 2 is increased autonomy and cost optimization. The model offers two tiers: Standard at $0.50 per 1M input / $2.50…
Google at May I/O 2026 has already begun "tightening the screws" and radically reshaping its infrastructure for developers. Gemini 3.5 Flash https://deepmind.google/models/gemini-3-5-flash/ The main "engine" of the announcement was the Gemini 3.5 Flash model, which precedes the upcoming 3.5 Pro. Google claims the model works significantly faster than previous generations and shows frontier-level…
A few interesting updates for May. Amidst the news about xAI, Anthropic also surprised many by announcing a partnership with SpaceX on May 6 to expand their computing power. Anthropic Discounts and Transition to New Pricing https://www.anthropic.com/news/higher-limits-spacex Anthropic announced a temporary "spring discount" on their API models. They also stopped blocking OpenClaw-style usage.…
Speaking of the major LLM players, xAI is the only one that hasn't been monetizing developers and programmers until now. It seems they are starting to change that. Cursor and xAI https://techsifted.com/posts/spacex-cursor-acquisition-april-2026/ SpaceX/xAI has secured an option to acquire Cursor for $60 billion. If the acquisition does not go through, Cursor will still receive $10 billion for…