Meituan introduced LongCat-2.0, a 1.6 trillion-parameter MoE model trained on domestic chips. The launch positions LongCat as a long-context agentic coding model, with official platform notes pointing to 1M-token context, native tool calling, and multi-step reasoning support. Read more
Anthropic launched Claude Sonnet 5 as its most performant Sonnet model yet. It is now the default model for Free and Pro plans, and is available in Claude Code and the API with lower-cost pricing aimed at Opus-like tool use, coding, and knowledge work. Read more
Google released Nano Banana 2 Lite and rolled Gemini Omni Flash to developers. Nano Banana 2 Lite targets four-second, low-cost image generation, while Omni Flash adds multimodal video generation and conversational editing in Google AI Studio and the Gemini API. Read more
Anthropic says Claude Fable 5 and Mythos 5 access has been restored after export controls were lifted. Fable 5 returns globally with a new safety classifier for a reported safeguard bypass, while Mythos 5 access remains tied to approved defensive-cyber partners. Read more
Vidu launched Vidu S1 for real-time interactive AI video characters. Users can create a character from an image, choose or clone a voice, and run live voice-driven calls with endless 540p video generation at 25fps or higher. Read more
Do you prefer to watch instead of read? Check out this video covering the top AI news this week:
InternScience released Agents-A1, a 35B MoE model for long-horizon agent work. It supports a served 256K context and is evaluated across search, engineering, scientific research, instruction following, and tool calling. Read more
ComfyUI introduced Comfy MCP in public beta so agents can operate Comfy workflows. The connector gives Claude, Codex, Cursor, and other agents access to image, video, 3D, and audio models plus ComfyUI workflow search, execution, and reuse. Read more
Meta released Brain2Qwerty v2 research for decoding words from non-invasive brain recordings. The system uses MEG brain signals and language-model context to reach 61% word accuracy overall and 78% for the best participant, while releasing training code and a dataset. Read more
NVIDIA GEAR introduced ASPIRE, a self-improving robotics system that writes and repairs code-as-policy programs. It inspects rollout traces, fixes failures, saves reusable robot skills, and improves tasks like two-arm handover from 20% to 92% in Robosuite. Read more
GLM is Z.AI’s flagship open-source model built for long, difficult engineering tasks like coding, debugging, and autonomous agent work. It delivers top-tier performance while being dramatically cheaper and faster than frontier closed models. Get 10% OFF HERE.
NVIDIA Isaac researchers introduced CHORD for learning dexterous robot manipulation from human demonstrations. Instead of only matching contact positions, CHORD matches the forces and torques that move objects, reaching 82.12% average success across 1,831 benchmark tasks and transferring to real robots. Read more
Perceptive BFM adapts human motion references to robot terrain in real time. A single policy follows raw motions like flips or fall recovery while using perception to choose footholds, clearance, posture, and contact timing for the terrain it sees. Read more
OmniContact turns contact-flow meta-skills into general humanoid loco-manipulation behaviors. Its dataset and live MuJoCo viewer show a humanoid carrying, pushing, kicking, recovering, and chaining tasks with closed-loop monitoring and VLM instruction support. Read more
LiveEdit brings diffusion-based video editing closer to real-time streaming. The ECCV 2026 project distills a bidirectional editing model into a causal streaming editor, using cache reuse to preserve backgrounds while reaching 12.66 FPS. Read more
Representation Distribution Matching trains a one-step image generator without an online teacher or adversary. Its iRDM approach matches generated and real feature distributions across frozen encoders, beating its four-step FLUX.2 teacher on GenEval and PickScore after about 90 H200 GPU-hours. Read more
Typeless is an intelligent AI voice dictation tool designed to turn your speech into polished, well-structured text in real time. It automatically cleans up your speech by removing filler words like “um,” fixing mid-sentence corrections, and formatting lists across various apps and devices. Try it for free today!
MrFlow speeds up text-to-image diffusion without training a new model. It samples cheaply at low resolution, upscales, re-encodes, and refines at high resolution, reporting more than 10x end-to-end speedup on Qwen-Image while preserving quality. Read more
PhysiFormer is a diffusion transformer for simulating 3D object mechanics in world coordinates. It predicts full mesh trajectories for rigid and elastic objects from initial positions, velocities, and materials, offering a learned alternative that can be faster than physics simulators after training. Read more
LUNA animates 3D human avatars without relying on traditional linear blend skinning. The ECCV 2026 project maps images, keypoints, sketches, and unseen characters into 3D Gaussian deformations, aiming for realistic zero-shot animation from 2D controls. Read more
ViDiHand uses video diffusion features to reconstruct smooth 4D hand motion. By fine-tuning a hand-aware branch while freezing the base diffusion transformer, it improves occlusion robustness, accuracy, and temporal smoothness across egocentric benchmarks. Read more
MuSViT is a foundation vision model built specifically for sheet music. It was pretrained on 9.7 million IMSLP pages and outperforms general-purpose encoders on music recognition, symbol detection, and score difficulty tasks. Read more
Meet Fish Audio, the most expressive AI voice model. Core features include text-to-speech, instant voice cloning, and a library of over 2 million voices. Create natural, emotionally rich AI voices built for your workflows. Try it for free.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.