RSS Amplifier

altsoph · Sep 8, 2025

My 6-month forecast for the AI industry

0
Sign in to vote or save

This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.

...remember this tweet...

As someone who worked as an analyst for many years, I have a guilty pleasure: I sometimes write predictions for myself about what will happen in a project/product/industry within a given time frame. Now, as the head of Research in Inworld.AI, I think about AI industry a lot. This exercise helps me later, when comparing the forecast with reality, to discover blind spots in my world model. I usually do this privately, but this time we had an interesting discussion about future of AI with a friend of mine, Fedor Zhdanov. Eventually, we decided to publish our forecasts separately, and compare them to reality later. Here is the Fedor’s forecast; and my projections for the AI-related industry for the next six months from the current moment are below. (BTW, here is my previous, initially private, forecast for the first half of 2025 + its review).

Extrapolating
source: xkcd:605

LLMs

  • harder and harder to compare different SotA models -- they have minimal differences on specific benchmarks/tasks (code/math/dialog/etc)

  • general reasoning is almost a commodity, but it is questionable how far it can boost quality in specific tasks (some additional gains would be via improved tool use, is there more?)

  • large MoE open-source models (mainly from China) are catching up or almost catching up with proprietary LLMs in quality

  • perf drift: more strange nuances (a specific provider affects quality through inference micro-optimization, implicit load routing, and invisible model updates that break usage patterns)

  • more cheaper / faster models, optimization towards mobile and in-browser use is on the way

  • jailbreaks for top models are still with us, including multimodal ones

  • longer contexts (but still with huge multi-span attention problems)

  • more experiments with diffusion text generation

  • more funny naming :)

Arch

  • no new architectures with any noticeable prospects for replacing transformers in production on the horizon for six months to a year (things of KAN and Mamba are still too exotic)

  • in specific tasks, transformers are maybe on par with diffusion networks — but they are still experimental, not product-ready yet.

  • even more micro-optimizations, see below.

Optimization

  • many micro-optimizations at the level of speculative inference/quantization/strange attention mods, CUDA kernels -- all those add some score in specific tasks, but sometimes break something else.

  • more local (on-device) models -- both open-source (like gemma-4, or small Chinese models) and proprietary (Apple Intelligence on-device model; Gemini Nano APIs), nothing fantastic but good for specific, well-defined tasks.

  • determinism budgeting: formalizing allowable nondeterminism/degradation in model routing/quantization/optimization to limit user-visible drift.

  • major players are still in search of market fit: what useful can be done by small and not-so-clever models

  • still no huge news on new proper consumer hardware (spec processors in mobile phones for end users)

Benchmarking

  • top1 @ OSWorld ~55%, livecodebench pro <3K, BALROG.nethack < 3%, BALROG.minihack < 30%.

  • preference leaderboards != benchmarks: human preferences are already differ from code/math benchmarks (bring me back my 4o) — the top 10 in different benchmarks are more mixed.

  • more hype based on targeted cherry-picking ("model X passed exams at Harvard better than all the meatbags and won all the olympiads")

  • (preference leaderboards + quality benchmarks) != stability/safety/reliability: no noticeable breakthroughs in reliability and stability in general (some individual benchmarks may fall, but the imbalance between good performance on pretentious tasks or competitions and complete nonsense on simple examples like famous 9.11/strawberry/etc remains)

  • the primary focus of efforts is on OS/browser/coding agents, but stability and predictability are still lacking.

Code generation

  • overall: the main frontier is stabilization, traceability, debugging, and productionalization of vibe-coding, with limited success so far. expect: sandboxing, audit trails, and deterministic safe modes.

  • global productionalization will be continued with focus on fast vibe-prototyping by solo startup founders :)

  • fully automated coding agents are still unstable and unreliable

  • new in-house solutions for businesses where keeping the code base proprietary is critical

  • still unknown how to make existing code generation models useful in the scope of a huge existing proprietary code base -> in-house benchmarking on specific tasks with metrics like task success, stuck-rate, rollback count, latency.

  • community is in search of intermediate abstractions (protocols, intermediate specs, agent-friendly logging) between code and human documentation to support more strict and controllable code generation

  • debugging of multi-agent systems is a nightmare; some experimental solutions may appear, but without significant success yet

  • maybe: human-time-to-verify as a cost metric.

Computer/browser-use agentic scenarios

  • still in a hype mode

  • more async assistance (deep research, scheduled tasks, proactive agents, etc).

  • more startups and solutions on in-house agent integration for businesses where data privacy is more critical than SotA quality

  • more integration attempts (ms office, google suite, adobe, cars, phones) -- all in search of profitable and useful patterns

  • prompt-injection & tool jailbreaks become more real risks as os-/browser-agents evolve

  • popular RAG frameworks are still unreliable for production without deep manual customization

  • popular multi-agent frameworks are still unreliable for production without deep manual customization

VLMs / MLLMs / Video gen (including diffusion)

  • more, better proprietary solutions

  • more open-source models from China

  • more impact on the multimedia creation market

  • quality issues and biases are still too frequent to make it fully unsupervised

  • optimistic: the first working video+LLM model to be demonstrated (or at least, announced) — if so, it could be OpenAI or Google

Training

  • new RL/xPO claims, but not solid -- likely important, still maturing

  • possible new optimizers

  • still no stable self-improving agents

  • still no lifelong training in practice

  • optimistic: MoE solutions with training special experts on domain / in-house data, to augment existing MoE models.

GPU / hardware

(not really my area of expertise)

  • major players are the same

  • optimistic: some announcements from Nvidia, AMD, and Apple, maybe; potentially something new from China (Huawei)

  • power and cooling water slowly become gating factors for new data-centers (but not a bottleneck yet); another potential future bottleneck is intra-datacenter networking

Robotics (anthropomorphic)

  • no considerable changes in 2025, but more funny videos and marketing press releases

  • no general household robots, some narrow traction (warehouse trials, pilot lines, but the economy does not add up yet)

  • main obstacles, again: cost, quality, absence of precise market fit (except military)

Self-driving

  • continues to infiltrate markets, slowly going closer to a commodity

  • self-flying is an important topic (mostly, military)

Other tech stuff

  • TTS/ASR is already a commodity; minor improvements, on-device migration, deeper customization (accents, emotions, etc), better/faster (1-shot) voice cloning.

  • potentially: real-time voice agents go mainstream, changing the call-center landscape -- limits: legal regulations.

  • Image gen (text2img): no big news, commodity, in-house productionalization for industry-specific tasks (textures, banners, assets, polygraphy, etc)

  • world models -- still in a hype mode, lack of stable results, lack of market fit.

  • NERF / Gaussian splatting / 3d generation -- similar, slightly less hype, more experiments, some incremental improvements.

  • music generation -- legal risks (like RIAA lawsuits), weak market fit, still several startups around.

  • Neurolinks -- still early experimental, years from mass market

  • drug discovery -- some experimental (newsworthy) results from labs like DM as a part of tech-PR, no AI-designed drug approvals yet.

  • other medicine applications -- regulations still block any significant progress.

Social impact

  • slightly outlined gradual shift from agi doom to ai bubble fears, but it is still a long way from a complete change of mood; no visible impact on the industry yet.

  • public dataset owners continue to commercialize their use in training (more dataset rights bundles), this changes the cost structure and legal risks, and increases the gap between well-capitalized labs and everyone else.

  • European vendors (like Mistral, PleIAs, etc) still have only limited success, but can benefit from local regulations

  • the "junior SWE" market is already significantly affected, making it harder to find a job as an IT engineer for recent graduates

  • successful agent/LLM usage in social networks, mainly in an abusive way, tightening of regulations / ToS / practices from the major networks

  • more impact on search traffic (AI answers on SERP, fewer clicks, fewer fact-checking), more impact on classic (textual) media

  • SEO -> GEO/AEO (Generative/Answer Engine Optimization): how to make a site so it pitches your data to incoming agents efficiently.

  • the majority of spam emails are AI-generated.

  • more attention on personalization for user (profile/history/cookies) -> privacy/targeting data market/GDPR-like things.

  • more low-quality content online (more AI-generated texts and videos) -> more signed content, things like C2PA, more importance of personal (human) recommendation institutes, closed communities, improved captchas, etc

  • more impact on education -- many experiments with unclear outcomes (too hard to make A/B there), classic education system disruption slowly continues

  • more security threats, including LLM-powered soceng and phishing large-scale attacks

  • more legal regulations around deepfakes

  • insurance products appear for AI failure cases

  • "vibecoding as ludomania" effects :)

Academia (NLP/ML/...) impact

  • a gap between industry and academia grows -> less weight of PhD (with specific exceptions)

  • conferences and online libraries are drowned in generated content; more automated reviews scandals -> academic community is in search for new forms (with no / limited success yet)

Political impact

  • more gov money in AI hype

  • more regulations

  • global market is more fragmented (soft/hard/wetware) to US/EU/China/others

  • too hard to predict :)


That’s all for today, folks!
Have ideas, corrections, or objections?
Welcome to the comments!

Want to read my postmortem for this forecast in 6 months? Subscribe to this blog to keep an eye on my updates:

Read on altsoph.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.