Kimi K3 is Moonshot AI's 2.8T-parameter open model with native vision and a 1-million-token context window. It is aimed at long coding, knowledge work, and reasoning; full weights are scheduled for release by July 27, 2026. Read more
Thinking Machines released Inkling, an open-weights multimodal model with 975B total parameters and 41B active parameters. The model supports up to 1M tokens and is built for customization through Tinker, including fine-tuning workflows that developers can run directly. Read more
OpenAI introduced GPT-Red, an internal automated red-teaming model trained to find prompt-injection failures. OpenAI says attacks generated by GPT-Red helped make GPT-5.6 much more robust while preserving normal capabilities. Read more
NVIDIA released Nemotron 3 Embed, a family of open embedding models for RAG, agent memory, code retrieval, and search. The 1B and 8B models are designed to improve retrieval quality while giving teams practical options for production deployment. Read more
Do you prefer to watch instead of read? Check out this video covering the top AI news this week:
OpenAI and Work Louder introduced Codex Micro, a $230 hardware controller for agentic work. Its keys, dial, joystick, and RGB status lights are meant to keep Codex actions and agent state visible while developers work. Read more
PrismML released Bonsai 27B, a compressed multimodal model based on Qwen3.6 27B that it says can run on a phone. The 1-bit version is 3.9GB and the ternary version is 5.9GB, which matters for private, local agents and offline use. Read more
Qualcomm AI Research introduced MobileWan, a 5B video diffusion system designed to run on commercial mobile hardware. It generates 5-second 480x832 videos at 16 FPS in about 20 seconds, narrowing the gap between mobile and server video models. Read more
Wan Streamer v0.3 turns real-time video generation into a 'world + event stream' problem. That lets the model keep a scene coherent while speech, motion, camera movement, and actions change at around 200 ms model-side latency. Read more
GLM is Z.AI's flagship open-source model built for long, difficult engineering tasks like coding, debugging, and autonomous agent work. It delivers top-tier performance while being dramatically cheaper and faster than frontier closed models. Get 10% OFF HERE.
Google DeepMind researchers showed that video generation backbones can be repurposed as general-purpose vision learners. GenCeption uses a text-steered feed-forward model for tasks like depth, segmentation, camera pose, and 4D keypoints with strong data efficiency. Read more
Schema is a new harness that reports near-ceiling scores on the ARC-AGI-3 public set without changing model weights. The authors say the results are public-set and self-reported, but the core idea is important: forcing models to build and test executable world models can matter as much as the model itself. Read more
Google released GNM, an open ecosystem for high-fidelity parametric human models, starting with GNM Head. It gives researchers and builders a permissively licensed 3D head model with controllable identity, expression, pose, eyeballs, teeth, and tongue. Read more
NVIDIA's ARDY generates controllable 3D human motion in real time from text and movement constraints. It is built for interactive animation, games, simulation, and humanoid robot control, where offline motion tools are too slow. Read more
Typeless is an intelligent AI voice dictation tool designed to turn your speech into polished, well-structured text in real time. It automatically cleans up your speech by removing filler words like "um," fixing mid-sentence corrections, and formatting lists across various apps and devices. Try it for free today!
Alibaba's Wan-Dancer generates minute-scale 720p dance videos that stay synchronized with music. The project uses global keyframe planning and local refinement to reduce the drift and identity changes that usually appear in long AI dance clips. Read more
Motion4Motion transfers movement between very different subjects without a shared skeleton or new training. It tracks dense motion flow from a source video and injects it into a frozen video diffusion model, enabling transfers like human-to-animal or a walking table demo. Read more
Mirelo and Kyutai open-sourced MuScriptor, an audio-to-MIDI model for full songs. It can separate a mix into MIDI tracks for voice, drums, bass, keys, and other instruments, which makes dense music easier to study or edit in a DAW. Read more
Lucida is an MIT-licensed background-removal model tuned for cases where common matting tools fail. It focuses on transparent objects, camouflage, text, glow effects, and illustrations, with weights on Hugging Face and a browser demo. Read more
Meet Fish Audio, the most expressive AI voice model. Core features include text-to-speech, instant voice cloning, and a library of over 2 million voices. Create natural, emotionally rich AI voices built for your workflows. Try it for free.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.