As someone who worked as an analyst for many years, I have a guilty pleasure: I sometimes write predictions for myself about what will happen in a project/product/industry within a given time frame. Now, as the head of Research in Inworld.AI, I think about AI industry a lot. This exercise helps me later, when comparing the forecast with reality, to discover blind spots in my world model. I usually do this privately, but this time we had an interesting discussion about future of AI with a friend of mine, Fedor Zhdanov. Eventually, we decided to publish our forecasts separately, and compare them to reality later. Here is the Fedor’s forecast; and my projections for the AI-related industry for the next six months from the current moment are below. (BTW, here is my previous, initially private, forecast for the first half of 2025 + its review).
LLMs
harder and harder to compare different SotA models -- they have minimal differences on specific benchmarks/tasks (code/math/dialog/etc)
general reasoning is almost a commodity, but it is questionable how far it can boost quality in specific tasks (some additional gains would be via improved tool use, is there more?)
large MoE open-source models (mainly from China) are catching up or almost catching up with proprietary LLMs in quality
perf drift: more strange nuances (a specific provider affects quality through inference micro-optimization, implicit load routing, and invisible model updates that break usage patterns)
more cheaper / faster models, optimization towards mobile and in-browser use is on the way
jailbreaks for top models are still with us, including multimodal ones
longer contexts (but still with huge multi-span attention problems)
more experiments with diffusion text generation
more funny naming :)
Arch
no new architectures with any noticeable prospects for replacing transformers in production on the horizon for six months to a year (things of KAN and Mamba are still too exotic)
in specific tasks, transformers are maybe on par with diffusion networks — but they are still experimental, not product-ready yet.
even more micro-optimizations, see below.
Optimization
many micro-optimizations at the level of speculative inference/quantization/strange attention mods, CUDA kernels -- all those add some score in specific tasks, but sometimes break something else.
more local (on-device) models -- both open-source (like gemma-4, or small Chinese models) and proprietary (Apple Intelligence on-device model; Gemini Nano APIs), nothing fantastic but good for specific, well-defined tasks.
determinism budgeting: formalizing allowable nondeterminism/degradation in model routing/quantization/optimization to limit user-visible drift.
major players are still in search of market fit: what useful can be done by small and not-so-clever models
still no huge news on new proper consumer hardware (spec processors in mobile phones for end users)
Benchmarking
top1 @ OSWorld ~55%, livecodebench pro <3K, BALROG.nethack < 3%, BALROG.minihack < 30%.
preference leaderboards != benchmarks: human preferences are already differ from code/math benchmarks (bring me back my 4o) — the top 10 in different benchmarks are more mixed.
more hype based on targeted cherry-picking ("model X passed exams at Harvard better than all the meatbags and won all the olympiads")
(preference leaderboards + quality benchmarks) != stability/safety/reliability: no noticeable breakthroughs in reliability and stability in general (some individual benchmarks may fall, but the imbalance between good performance on pretentious tasks or competitions and complete nonsense on simple examples like famous 9.11/strawberry/etc remains)
the primary focus of efforts is on OS/browser/coding agents, but stability and predictability are still lacking.
Code generation
overall: the main frontier is stabilization, traceability, debugging, and productionalization of vibe-coding, with limited success so far. expect: sandboxing, audit trails, and deterministic safe modes.
global productionalization will be continued with focus on fast vibe-prototyping by solo startup founders :)
fully automated coding agents are still unstable and unreliable
new in-house solutions for businesses where keeping the code base proprietary is critical
still unknown how to make existing code generation models useful in the scope of a huge existing proprietary code base -> in-house benchmarking on specific tasks with metrics like task success, stuck-rate, rollback count, latency.
community is in search of intermediate abstractions (protocols, intermediate specs, agent-friendly logging) between code and human documentation to support more strict and controllable code generation
debugging of multi-agent systems is a nightmare; some experimental solutions may appear, but without significant success yet
maybe: human-time-to-verify as a cost metric.
Computer/browser-use agentic scenarios
still in a hype mode
more async assistance (deep research, scheduled tasks, proactive agents, etc).
more startups and solutions on in-house agent integration for businesses where data privacy is more critical than SotA quality
more integration attempts (ms office, google suite, adobe, cars, phones) -- all in search of profitable and useful patterns
prompt-injection & tool jailbreaks become more real risks as os-/browser-agents evolve
popular RAG frameworks are still unreliable for production without deep manual customization
popular multi-agent frameworks are still unreliable for production without deep manual customization
VLMs / MLLMs / Video gen (including diffusion)
more, better proprietary solutions
more open-source models from China
more impact on the multimedia creation market
quality issues and biases are still too frequent to make it fully unsupervised
optimistic: the first working video+LLM model to be demonstrated (or at least, announced) — if so, it could be OpenAI or Google
Training
new RL/xPO claims, but not solid -- likely important, still maturing
possible new optimizers
still no stable self-improving agents
still no lifelong training in practice
optimistic: MoE solutions with training special experts on domain / in-house data, to augment existing MoE models.
GPU / hardware
(not really my area of expertise)
major players are the same
optimistic: some announcements from Nvidia, AMD, and Apple, maybe; potentially something new from China (Huawei)
power and cooling water slowly become gating factors for new data-centers (but not a bottleneck yet); another potential future bottleneck is intra-datacenter networking
Robotics (anthropomorphic)
no considerable changes in 2025, but more funny videos and marketing press releases
no general household robots, some narrow traction (warehouse trials, pilot lines, but the economy does not add up yet)
main obstacles, again: cost, quality, absence of precise market fit (except military)
Self-driving
continues to infiltrate markets, slowly going closer to a commodity
self-flying is an important topic (mostly, military)
Other tech stuff
TTS/ASR is already a commodity; minor improvements, on-device migration, deeper customization (accents, emotions, etc), better/faster (1-shot) voice cloning.
potentially: real-time voice agents go mainstream, changing the call-center landscape -- limits: legal regulations.
Image gen (text2img): no big news, commodity, in-house productionalization for industry-specific tasks (textures, banners, assets, polygraphy, etc)
world models -- still in a hype mode, lack of stable results, lack of market fit.
NERF / Gaussian splatting / 3d generation -- similar, slightly less hype, more experiments, some incremental improvements.
music generation -- legal risks (like RIAA lawsuits), weak market fit, still several startups around.
Neurolinks -- still early experimental, years from mass market
drug discovery -- some experimental (newsworthy) results from labs like DM as a part of tech-PR, no AI-designed drug approvals yet.
other medicine applications -- regulations still block any significant progress.
Social impact
slightly outlined gradual shift from agi doom to ai bubble fears, but it is still a long way from a complete change of mood; no visible impact on the industry yet.
public dataset owners continue to commercialize their use in training (more dataset rights bundles), this changes the cost structure and legal risks, and increases the gap between well-capitalized labs and everyone else.
European vendors (like Mistral, PleIAs, etc) still have only limited success, but can benefit from local regulations
the "junior SWE" market is already significantly affected, making it harder to find a job as an IT engineer for recent graduates
successful agent/LLM usage in social networks, mainly in an abusive way, tightening of regulations / ToS / practices from the major networks
more impact on search traffic (AI answers on SERP, fewer clicks, fewer fact-checking), more impact on classic (textual) media
SEO -> GEO/AEO (Generative/Answer Engine Optimization): how to make a site so it pitches your data to incoming agents efficiently.
the majority of spam emails are AI-generated.
more attention on personalization for user (profile/history/cookies) -> privacy/targeting data market/GDPR-like things.
more low-quality content online (more AI-generated texts and videos) -> more signed content, things like C2PA, more importance of personal (human) recommendation institutes, closed communities, improved captchas, etc
more impact on education -- many experiments with unclear outcomes (too hard to make A/B there), classic education system disruption slowly continues
more security threats, including LLM-powered soceng and phishing large-scale attacks
more legal regulations around deepfakes
insurance products appear for AI failure cases
"vibecoding as ludomania" effects :)
Academia (NLP/ML/...) impact
a gap between industry and academia grows -> less weight of PhD (with specific exceptions)
conferences and online libraries are drowned in generated content; more automated reviews scandals -> academic community is in search for new forms (with no / limited success yet)
Political impact
more gov money in AI hype
more regulations
global market is more fragmented (soft/hard/wetware) to US/EU/China/others
too hard to predict :)
That’s all for today, folks!
Have ideas, corrections, or objections?
Welcome to the comments!
Want to read my postmortem for this forecast in 6 months? Subscribe to this blog to keep an eye on my updates:

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.