DeepSeek releases DeepSeek-V4-Flash-0731 on Hugging Face under an MIT license. The 304B-parameter text model matches the intelligence of GLM 5.2 and Opus 4.8, but is 100x cheaper and much smaller, presenting a new frontier in efficiency. Read more
Google DeepMind introduced Gemini Robotics 2, a set of models that help robots reason, move, and work together in the real world. It can control full humanoids from feet to fingertips, run on-device, and adapt to new robot bodies with only a few hours of data. Read more
OpenAI released Codex Security as a CLI and TypeScript SDK for finding, validating, and fixing security bugs in code. Developers can run scans locally or in CI, compare scan histories, and use either ChatGPT sign-in or API-key based authentication. Read more
Do you prefer to watch instead of read? Check out this video covering the top AI news this week:
Thinking Machines Lab released Inkling-Small, an open-weights multimodal reasoning model that is much smaller than its flagship Inkling model. It has 276B total parameters with 12B active, supports images and audio, reaches up to a 1M-token context window, and is available through Hugging Face and Tinker. Read more
AMD introduced Instella-MoE, a fully open mixture-of-experts language model trained on AMD Instinct GPUs with ROCm. The release includes model weights from each training stage, data mixtures, code, and a 16B-total, 2.8B-active architecture meant to show that open AMD-based training stacks can compete. Read more
Google added natural voice control to the Gemini app for macOS. Long-pressing the Fn key lets users dictate polished text into any app, and optional Gemini reasoning can summarize files, rewrite selected text, or edit images based on what is on screen. Read more
Hugging Face published a technical timeline of a July 2026 agent intrusion that produced about 17,600 actions over several days. The write-up explains how the agent chained ordinary security weaknesses across trust boundaries, and why defenders need tighter isolation, short-lived credentials, and better cross-system detection. Read more
Unlock next-level video creation with MiniMax H3, the powerful new model that accepts text, image, audio, and video as references through its Omni Reference system. Generate, edit, and transform content with unprecedented flexibility. Deliver native 2K video up to 15 seconds, purpose-built for film, advertising, business, and more. Try it today!
Vivix-A1 is a real-time model for full-body AI characters that can listen, speak, move, and act inside open worlds. It supports voice, text, and image inputs during a live interaction, so a character can be redirected while it is already talking or moving. Read more
ByteDance Seed launched Seedance 2.5, a new audio-video generation model built for longer 30-second stories. It gives creators tighter reference control, smoother motion, and editing tools like green-screen and white-model control for more production-style workflows. Read more
BeingBeyond introduced Being-H0.8, a tactile-aware embodied foundation model for robots. It learns from more than 500,000 hours of egocentric human video, adds pseudo-touch labels through TactoHand, and uses a shared hand/action representation for human hands, robot hands, and grippers. Read more
Wonder is a video world model that turns an image or video into an interactive world users can navigate. It keeps memory of previous views, reveals unseen areas, and supports up to one-minute rollouts at constant latency with 16 FPS generation. Read more
Nyra Labs released CrisperWhisper 2.0 for verbatim speech recognition. The model can switch between word-for-word transcripts and cleaned-up intended transcripts, supports multiple languages, and reports precise word-level timing for speech that includes fillers, stutters, and vocal sounds. Read more
PhiZero proposes a world model built around a learned physical language instead of plain natural language. It predicts future scene changes as compact transition tokens first, then renders them into video, which could help models reason about physics before making pixels. Read more
Typeless is an intelligent AI voice dictation tool designed to turn your speech into polished, well-structured text in real time. It automatically cleans up your speech by removing filler words like “um,” fixing mid-sentence corrections, and formatting lists across various apps and devices. Try it for free today!
ID-V2V is a Netflix and Eyeline Labs research project for restyling videos while preserving a person’s identity and performance. It uses relit face regions, facial normals, depth, and edited keyframes so style changes can spread through a clip without losing expressions, gaze, or lip sync. Read more
PRISM is a robot-control method that learns useful physical cues from signals a robot already has. By modeling polynomial interactions in proprioception, it improves humanoid survival, manipulation success, and long-horizon LIBERO tasks without adding new sensors. Read more
ReDesign turns a flat design screenshot back into editable design structure. The ECCV 2026 project decomposes an image into text, shapes, images, groups, and layer order, then exports an editable JSON hierarchy for recovering lost design files. Read more
ShadowDancer teaches video world models to replay actions from demonstrations instead of labels or text commands. It learns from paired videos that share the same dynamics but change the scene, character, or lighting, then transfers those actions into new worlds and longer rollouts. Read more
Meet Fish Audio, the most expressive AI voice model. Core features include text-to-speech, instant voice cloning, and a library of over 2 million voices. Create natural, emotionally rich AI voices built for your workflows. Try it for free.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.