RSS Amplifier

AG+ (AI Daily News) · Aug 7, 2026

AMD Just Bought One of AI’s Fastest Inference Startups

0
Sign in to vote or save

AJ Green · AG+ (AI Daily News)

AMD just acquired a startup that came out of stealth a few months ago claiming it can run an AI model at roughly 17,000 tokens per second for a single user.

There is a catch: Taalas is getting those speeds by physically embedding a specific model into silicon, and its first chip runs Llama 3.1 8B — a much smaller, older model than the frontier systems we use today. But that is exactly why this acquisition is interesting. Taalas is attacking one of AI infrastructure’s most expensive problems - constantly moving enormous model weights between memory and compute - and believes specialized silicon can dramatically reduce the hardware, energy, and cost required to serve intelligence.

In today’s AI news:

  • AMD acquires Taalas and its 17,000 token-per-second AI chips

  • AI starts designing complete biological systems

  • Anthropic rethinks Fable 5’s biology guardrails

  • DeepSeek signals the end of AI’s race to the bottom

  • Today’s Top Tools + Quick News

News: AMD has acquired Toronto-based Taalas, an AI chip startup building custom silicon that physically embeds model weights into the chip itself. Taalas emerged from stealth earlier this year with its HC1 technology demonstrator running Llama 3.1 8B at a claimed 17,000 tokens per second per user, and AMD now plans to bring the company’s technology into its broader AI accelerator roadmap.

Details:

  • Taalas takes a fundamentally different approach from conventional GPUs: its Hardcore Model architecture stores model weights directly in silicon, attacking the expensive movement of data between memory and compute.

  • Its first HC1 demonstrator is a 53-billion-transistor, 6nm chip running Llama 3.1 8B, with Taalas reporting roughly 17,000 tokens per second for a single user.

  • That number needs context: Llama 3.1 8B is no longer a frontier model, and the 17K tok/s result comes from Taalas’ own testing. The demonstration proves extraordinary speed on a specialized workload, not that HC1 delivers frontier-model intelligence.

  • The architecture’s potential advantage is efficiency. Taalas says it can remove the need for technologies such as HBM, advanced packaging, 3D stacking, and liquid cooling, potentially lowering power consumption and inference cost.

  • Taalas had already raised roughly $219M before the acquisition, including $169M earlier this year, and AMD says the technology will sit alongside - rather than simply replace - its Instinct GPU strategy.

Why It Matters: The interesting thing about Taalas is not simply 17,000 tokens per second - it is the tradeoff the company is willing to make to get there. GPUs are incredibly flexible, but that flexibility comes with a massive memory and energy tax; Taalas sacrifices some of it by physically optimizing silicon around specific models. That could become incredibly compelling in a world where a handful of stable models are serving billions of calls from people and agents every day. The risk is that models are evolving extraordinarily quickly: if architectures, weights, or techniques change faster than Taalas can design and manufacture new silicon, today’s optimization can become tomorrow’s constraint.

News: Stanford and Arc Institute researchers used genome language models to design complete bacteriophage genomes, synthesized hundreds of the designs, and found that 16 successfully propagated and inhibited their intended E. coli strains. Arc describes the work as its first demonstration of AI-designed and experimentally validated organisms — a very different frontier from systems such as AlphaFold, which predict biological structure rather than generate complete genomes that can be physically built and tested.

Details:

  • Researchers used Evo 1 and Evo 2 to generate variants of ΦX174, a small, well-studied bacteriophage with only 11 genes.

  • Of 285 tested designs, 16 worked, successfully propagating and inhibiting the appropriate bacterial strains without affecting unrelated strains.

  • This builds on a much broader revolution in computational biology. AlphaFold 3 showed how AI can predict the structures and interactions of proteins, DNA, RNA, ligands, and other biomolecules; genome models like Evo are pushing toward designing new biological systems.

  • Arc says the generated phages retained host specificity, which matters if systems like these eventually contribute to treatments targeting antibiotic-resistant bacteria.

  • The experiment is still relatively constrained — ΦX174 has a tiny genome compared with more complex organisms — but it is a compelling proof of concept for genome-scale generative biology.

Why It Matters: This is the version of AI that people around the world are genuinely hopeful about. Over the last few weeks, I’ve been traveling with Anthropic and asking people in different cities what they hope AI ultimately does for humanity - and scientific advancement comes up everywhere. AlphaFold showed us what happens when AI becomes extraordinarily good at understanding biology; work like this asks what happens when AI gets good at designing it. That upside is profound, but so is the dual-use risk. If there are two capability domains where we should demand exceptional safeguards as AI becomes more powerful, they are biology and cybersecurity.

News: Anthropic launched Claude Fable 5 with deliberately conservative safeguards that route many biology, chemistry, and other sensitive requests away from the company’s most capable model. After users and researchers showed just how broadly those restrictions were firing, Anthropic has acknowledged that the biology classifiers are broader than it would like and says it is working to make them more precise.

Details:

  • Anthropic explicitly lists the majority of biology, chemistry, and life-sciences queries among areas where users may currently encounter fallbacks, including some benign healthcare, diagnostics, education, and biotech work.

  • Independent biomedical research found refusal rates ranging from 8% to 99.4% depending on the benchmark, while Fable’s performance was extremely strong on questions it actually answered.

  • Anthropic says these blocking safeguards are intentionally broad and that it is continuing to improve them so they target risky behavior more precisely without unnecessarily degrading legitimate use.

  • This sits alongside similar work in cybersecurity, where Anthropic is publicly detailing the classifier systems and jailbreak frameworks used to control access to advanced capabilities.

Why It Matters: This is the fine line every frontier lab is going to have to walk. As models become genuinely useful for biology, cybersecurity, and advanced scientific work, giving everyone unrestricted access to every capability is difficult to justify - but blocking harmless and beneficial work defeats much of the reason we are building these systems in the first place. AI safety at the frontier increasingly is not a binary question of whether a model can do something; it is a question of who should be able to do what, under which conditions, and with how much friction... and ultimately who gets to decide those parameters.

News: DeepSeek has warned API customers that it plans to significantly increase pricing, saying overall prices will rise “by a relatively large margin” while leaving the exact pricing structure and timing for a later announcement. After becoming one of the companies most responsible for resetting expectations around how cheaply high-performing AI could be served, the reversal may be an early signal that the industry is moving from buying adoption with price to monetizing outcomes.

Details:

  • DeepSeek has not yet disclosed the exact increase or effective date, and it has not publicly said that rising compute costs are the reason - that explanation remains an inference rather than an official rationale.

  • The move is particularly interesting because DeepSeek has spent much of the last year pulling prices in the opposite direction, including cutting prices again earlier this year.

  • DeepSeek also has a unique constraint: because its models are open weight, customers can potentially self-host or move to another inference provider running the same model if the official API becomes too expensive.

  • The most important thing to watch is what comes with the increase: a new flagship model, tighter free allowances, enterprise tiers, agent-oriented pricing, or other changes would make this look much more like a broader commercialization shift.

Why It Matters: For years, we have talked about AI economics in dollars per million tokens. I increasingly think that is the wrong unit. When a model is writing software, conducting research, operating agents, or completing an entire business process, the buyer cares much more about what the completed outcome costs than how many tokens were consumed along the way. DeepSeek helped accelerate the race toward cheaper intelligence; this price increase may be one signal that the next competition is about proving the intelligence is valuable enough to charge for.

  • Wan-Streamer — Alibaba researchers’ real-time audio-visual model that treats video as a persistent “world + event stream,” pushing toward interactive video generation rather than one-shot clips.

  • TranslateGemma — Google’s open translation models built on Gemma, designed to bring high-quality multilingual translation across local devices and larger compute environments.

  • Muse Code — Meta’s new AI coding agent, entering direct competition with tools like Claude Code and Codex as coding becomes one of the most commercially important agent categories.

  • OpenAI asked a federal judge to dismiss Apple’s trade-secret lawsuit, arguing that Apple has failed to sufficiently identify the trade secrets allegedly taken or demonstrate that OpenAI misappropriated them. The dispute is becoming another front in the competition to control AI talent, devices, and the post-smartphone computing layer.

  • Meta’s Muse Spark 1.1 accessed another company’s systems during a cybersecurity evaluation after a testing misconfiguration gave the model unintended internet access. Investigators emphasized that this was not a sophisticated sandbox escape, but the incident reinforces something increasingly important: agent safety depends just as much on the infrastructure and permissions surrounding the model as the model itself.

  • OpenAI is giving Free and Go users unlimited everyday text chats alongside a new Think option for harder questions, while GPT-5.6 Luna becomes the default model for those users. This continues the extraordinary compression in the cost of accessing capable intelligence.

  • Google Maps’ Ask Maps lets users ask Gemini complex questions about where to go, what to do, and how to navigate real-world decisions directly inside Maps. It is another step toward search products becoming conversational decision layers rather than directories that simply return places.

Read the original on ajsai.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.