RSS Amplifier

AI Search · Jul 12, 2026

GPT 5.6, Grok 4.5, GPT Live, LingBot World 2, Seedream 5, Muse Spark: AI NEWS

0
Sign in to vote or save

AI Search · AI Search

Robbyant’s LingBot-World 2.0 is an open world model for exploring persistent, interactive environments in real time. It targets 720p at 60 fps with sub-second control latency, supports collaborative steering and event-driven interactions, and points toward robot simulation beyond short video clips. Read more

OpenAI is rolling out GPT‑Live, a full-duplex voice model that can listen and speak at the same time for more natural conversation. It keeps talking while delegating search and hard reasoning to frontier models, and is rolling out globally in ChatGPT before an API release. Read more

OpenAI released GPT-5.6, a new family of smarter and more efficient AI models including the flagship Sol. It delivers top performance in coding, science, knowledge work, and complex tasks while using fewer tokens and costing less than before — with special “ultra” mode that teams up multiple agents for really hard jobs. Smaller models like Terra and Luna make advanced AI more affordable for everyday use. Read more

Meta’s Muse Spark 1.1 is a multimodal reasoning model built for agents, computer use, coding, and long-running workflows. Its 1-million-token context and parallel subagent orchestration are now available through a public-preview Meta Model API. Read more

Do you prefer to watch instead of read? Check out this video covering the top AI news this week:

xAI’s Grok 4.5 is a model aimed at coding, agentic tasks, and knowledge work, with a focus on engineering performance and efficiency. It is available in Grok Build, Cursor, and the xAI console, where xAI says it serves at 80 tokens per second and costs $2 per million input tokens and $6 per million output tokens. Read more

Meta has launched Muse Image and previewed Muse Video as its first media-generation models from Meta Superintelligence Labs. Muse Image can search, write code, self-refine, edit and combine references, while Muse Video targets high-fidelity video with native audio. Read more

ByteDance’s Seedream 5.0 Pro is a multimodal image generator built for advanced reasoning and professional visual production. It focuses on text-rich infographics, sketch- and annotation-guided editing with layer separation, photorealistic lighting and skin, and native support for a dozen languages. Read more

ABot-World turns a single RTX 5090 desktop GPU into an interactive world simulator that responds to user actions. The project reports 720p generation at 16 fps with 1.2-second latency and 19 GB of memory, enabling open-ended rollouts instead of fixed-length videos. Read more

GLM is Z.AI’s flagship open-source model built for long, difficult engineering tasks like coding, debugging, and autonomous agent work. It delivers top-tier performance while being dramatically cheaper and faster than frontier closed models. Get 10% OFF HERE.

AlayaWorld is a 15-billion-parameter world model that generates playable video worlds at 720p and 24 fps for more than a minute. It combines a 3D scene cache with frame-history memory so users can move the camera, switch prompts mid-scene, and revisit places without the world changing as much. Read more

MIRA is a 5-billion-parameter model that simulates four-player Rocket League-style matches directly from video and player actions. Built by General Intuition and Kyutai with Epic Games, it runs a shared 2v2 world at 20 fps without a traditional game or physics engine, and the team is releasing code and data. Read more

Wan Streamer v0.2 raises real-time generated video from 192×336 to 640×368 while keeping about 200 ms of model-side latency. Its thinker–performer setup uses one GPU for perception and state updates plus parallel GPUs for the higher-resolution video, making scene-grounded mid-shot agents possible at 25 fps. Read more

Tencent’s open-source Hy3 is a smaller agent model tuned for reasoning, coding, office work, long contexts, and reliable tool use. Tencent reports a 2.67/4 score from a 270-expert blind evaluation, lower hallucination and multi-turn error rates than its preview, and Apache 2.0 weights with API pricing starting at 1 RMB per million input tokens. Read more

Typeless is an intelligent AI voice dictation tool designed to turn your speech into polished, well-structured text in real time. It automatically cleans up your speech by removing filler words like “um,” fixing mid-sentence corrections, and formatting lists across various apps and devices. Try it for free today!

ProxyPose turns one clicked pixel in a monocular video into a full 6-DoF motion trajectory by first generating a proxy video. That simple interface lets it track rigid objects, cameras, faces, event-camera footage, and single-photon video, though fast motion and reflective or textureless surfaces can still cause drift. Read more

PixWorld unifies 3D scene reconstruction and generation in one pixel-space diffusion model. Instead of hiding the scene inside a latent code, it trains directly on rendered views and adds geometry-aware supervision, aiming for more faithful 3D structure with fewer representation-loss problems. Read more

SeFi-Image is a text-to-image foundation model that denoises semantic structure slightly ahead of texture details. Its 1B, 2B, and 5B versions use that two-stream approach to improve layout, long text, and visual fidelity while the 5B model was trained with 125,000 A800 GPU hours. Read more

A study evaluates a humanoid robot for laparoscopic surgery using teleoperation, benchtop tests, user studies, and live porcine experiments. The work shows that general-purpose humanoid hardware can perform surgical tasks, but it also highlights the precision, control, and safety gaps that must be closed before clinical use. Read more

Meet Fish Audio, the most expressive AI voice model. Core features include text-to-speech, instant voice cloning, and a library of over 2 million voices. Create natural, emotionally rich AI voices built for your workflows. Try it for free.

Read the original on aisearch.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.