"{"canSubmit":false,"dailyPapers":[{"paper":{"id":"2608.09888","authors":[{"_id":"6a7af16fb7183653340c11cd","user":{"_id":"66d17d0d2fa8a088dc6992b3","avatarUrl":"/avatars/0aeb72e86b49d85db5577b507751147a.svg","isPro":false,"fullname":"Björn Engdahl","user":"bjornengdahl","type":"user","name":"bjornengdahl"},"name":"Björn Engdahl","status":"claimed_verified","statusLastChangedAt":"2026-08-11T16:45:04.704Z","hidden":false},{"_id":"6a7af16fb7183653340c11ce","user":{"_id":"63a48fe7658851481f766a24","avatarUrl":"/avatars/4d1738e80fd7677a81197fb158ba8c44.svg","isPro":false,"fullname":"Adrian Kosowski","user":"dxtrous","type":"user","name":"dxtrous"},"name":"Adrian Kosowski","status":"claimed_verified","statusLastChangedAt":"2026-08-12T00:45:04.441Z","hidden":false},{"_id":"6a7af16fb7183653340c11cf","user":{"_id":"651310e3b7994cff61d71320","avatarUrl":"/avatars/161cb412c7cb284fa7483e790807e4d3.svg","isPro":false,"fullname":"Jan Chorowski","user":"janchorowski","type":"user","name":"janchorowski"},"name":"Jan Chorowski","status":"claimed_verified","statusLastChangedAt":"2026-08-11T16:45:04.695Z","hidden":false},{"_id":"6a7af16fb7183653340c11d0","user":{"_id":"68dd50bb9d2cc46fef645799","avatarUrl":"/avatars/699f84aa771851854a4201cb1df3543b.svg","isPro":false,"fullname":"Zuzanna Stamirowska","user":"Brain-synapse","type":"user","name":"Brain-synapse"},"name":"Zuzanna Stamirowska","status":"claimed_verified","statusLastChangedAt":"2026-08-11T16:45:04.700Z","hidden":false},{"_id":"6a7af16fb7183653340c11d1","name":"Przemysław Uznański","hidden":false},{"_id":"6a7af16fb7183653340c11d2","name":"Junlin Jiang","hidden":false},{"_id":"6a7af16fb7183653340c11d3","user":{"_id":"63924fd62bfa48413aa0caec","avatarUrl":"/avatars/9ebf245b9ca187a213a2b15f7853ea08.svg","isPro":false,"fullname":"Rohan Phadke","user":"rohanphadke","type":"user","name":"rohanphadke"},"name":"Rohan Phadke","status":"claimed_verified","statusLastChangedAt":"2026-08-12T00:45:04.516Z","hidden":false},{"_id":"6a7af16fb7183653340c11d4","name":"Remigiusz Kinas","hidden":false},{"_id":"6a7af16fb7183653340c11d5","name":"Richard Zhong","hidden":false}],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"BDH-CQ: In-Context Learning with Recurrent Latent Reasoning","submittedOnDailyBy":{"_id":"651310e3b7994cff61d71320","avatarUrl":"/avatars/161cb412c7cb284fa7483e790807e4d3.svg","isPro":false,"fullname":"Jan Chorowski","user":"janchorowski","type":"user","name":"janchorowski"},"summary":"We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning. We evaluate the model on the public ARC-AGI-1 evaluation set and use controlled ARC-like interventions to study what it learns from demonstrations, how consistently it applies an inferred transformation, and which concepts remain difficult. A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of \\$0.0007 per task. This operating point breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency.","upvotes":639,"discussionId":"6a7af170b7183653340c11d6","projectPage":"https://pathway.com/blog/pathway-150m-model-breaks-arc-agi-1-cost-efficiency-frontier","githubRepo":"https://github.com/pathwaycom/arc-task-gen","githubRepoAddedBy":"user","ai_summary":"A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1.","ai_keywords":["in-context learning","recurrent latent reasoning","high-dimensional latent space","ARC-AGI-1","cost-accuracy Pareto frontier"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3227,"organization":{"_id":"6697b1a8b981e21b48bc575f","name":"pathwaycom","fullname":"Pathway","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6697aead07b36ccd016f0888/nOEvf0HpKBQJcVDfljgAc.png"}},"publishedAt":"2026-08-09T20:00:00.000Z","title":"BDH-CQ: In-Context Learning with Recurrent Latent Reasoning","summary":"We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning. We evaluate the model on the public ARC-AGI-1 evaluation set and use controlled ARC-like interventions to study what it learns from demonstrations, how consistently it applies an inferred transformation, and which concepts remain difficult. A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of \\$0.0007 per task. This operating point breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2608.09888.png","numComments":3,"upvoted":false,"submittedBy":{"_id":"651310e3b7994cff61d71320","avatarUrl":"/avatars/161cb412c7cb284fa7483e790807e4d3.svg","fullname":"Jan Chorowski","name":"janchorowski","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":14,"isUserFollowing":false},"organization":{"_id":"6697b1a8b981e21b48bc575f","name":"pathwaycom","fullname":"Pathway","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6697aead07b36ccd016f0888/nOEvf0HpKBQJcVDfljgAc.png"},"isAuthorParticipating":true},{"paper":{"id":"2605.31264","authors":[{"_id":"6a1cf553808ddbc3c7d434af","name":"Tianyi Zhou","hidden":false},{"_id":"6a1cf553808ddbc3c7d434b0","name":"Dongrui Liu","hidden":false},{"_id":"6a1cf553808ddbc3c7d434b1","name":"Leitao Yuan","hidden":false},{"_id":"6a1cf553808ddbc3c7d434b2","name":"Jing Shao","hidden":false},{"_id":"6a1cf553808ddbc3c7d434b3","name":"Xia Hu","hidden":false}],"publishedAt":"2026-05-29T00:00:00.000Z","submittedOnDailyAt":"2026-06-01T00:00:00.000Z","title":"COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation","submittedOnDailyBy":{"_id":"66e2624a436a1798365e4581","avatarUrl":"/avatars/6c605807d34faa8fb505e135a4b47776.svg","isPro":false,"fullname":"Qihan Ren","user":"jasonrqh","type":"user","name":"jasonrqh"},"summary":"LLM agents are increasingly expected not only to complete isolated tasks, but also to carry bounded representations of human expertise, judgment, and interaction style. Building such person-grounded agents remains difficult because actionable knowledge associated with a person or role is usually embedded in heterogeneous traces rather than written as clean instructions. Existing memory and persona systems capture fragments of this evidence, while skill frameworks provide portable packaging formats; however, there is no end-to-end workflow for distilling these traces into inspectable, correctable, and agent-usable skills. We present an automated trace-to-skill distillation system for generating person-grounded AI skills via expert knowledge distillation. Given materials from a target person or role, COLLEAGUE.SKILL produces a versioned skill package with two coordinated tracks: a capability track for practices, mental models, and decision heuristics, and a bounded behavior track for communication style, interaction rules, and correction history. The package can be inspected, invoked, updated through natural-language feedback, rolled back, installed across agent hosts, and optionally prepared for controlled distribution. We describe the artifact contract, generation workflow, correction lifecycle, deployment surface, and domain presets implemented in the open-source system. At the time of writing, the public repository has approximately 18.5k GitHub stars; the gallery lists 215 skills from 165 contributors and more than 100k cumulative stars across listed skill cards. The system illustrates how person-grounded skills can be represented as portable, correctable packages rather than opaque prompts or hidden memories.","upvotes":126,"discussionId":"6a1cf554808ddbc3c7d434b4","githubRepo":"https://github.com/titanwings/colleague-skill","githubRepoAddedBy":"user","ai_summary":"Person-grounded AI skills are automatically distilled from heterogeneous traces into inspectable, correctable packages that capture both capabilities and behavioral patterns.","ai_keywords":["expert knowledge distillation","person-grounded agents","trace-to-skill distillation","capability track","behavior track","skill package","correction lifecycle","deployment surface","domain presets"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":23074,"organization":{"_id":"6747ee5decec679eafb90450","name":"ShanghaiAiLab","fullname":"shanghai ailab ","avatar":"https://www.gravatar.com/avatar/6cd2acf412ad103653d9ce14a1aacc19?d=retro&size=100"}},"publishedAt":"2026-05-28T20:00:00.000Z","title":"COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation","summary":"LLM agents are increasingly expected not only to complete isolated tasks, but also to carry bounded representations of human expertise, judgment, and interaction style. Building such person-grounded agents remains difficult because actionable knowledge associated with a person or role is usually embedded in heterogeneous traces rather than written as clean instructions. Existing memory and persona systems capture fragments of this evidence, while skill frameworks provide portable packaging formats; however, there is no end-to-end workflow for distilling these traces into inspectable, correctable, and agent-usable skills. We present an automated trace-to-skill distillation system for generating person-grounded AI skills via expert knowledge distillation. Given materials from a target person or role, COLLEAGUE.SKILL produces a versioned skill package with two coordinated tracks: a capability track for practices, mental models, and decision heuristics, and a bounded behavior track for communication style, interaction rules, and correction history. The package can be inspected, invoked, updated through natural-language feedback, rolled back, installed across agent hosts, and optionally prepared for controlled distribution. We describe the artifact contract, generation workflow, correction lifecycle, deployment surface, and domain presets implemented in the open-source system. At the time of writing, the public repository has approximately 18.5k GitHub stars; the gallery lists 215 skills from 165 contributors and more than 100k cumulative stars across listed skill cards. The system illustrates how person-grounded skills can be represented as portable, correctable packages rather than opaque prompts or hidden memories.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2605.31264.png","numComments":3,"upvoted":false,"submittedBy":{"_id":"66e2624a436a1798365e4581","avatarUrl":"/avatars/6c605807d34faa8fb505e135a4b47776.svg","fullname":"Qihan Ren","name":"jasonrqh","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false},"organization":{"_id":"6747ee5decec679eafb90450","name":"ShanghaiAiLab","fullname":"shanghai ailab ","avatar":"https://www.gravatar.com/avatar/6cd2acf412ad103653d9ce14a1aacc19?d=retro&size=100"},"isAuthorParticipating":false},{"paper":{"id":"2412.20138","authors":[{"_id":"6847673b3ec10bdd8ab4dd5e","user":{"_id":"619b6a24030b707ff1f05dee","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/619b6a24030b707ff1f05dee/8FZjePfOuTtN3uvTmaEg5.jpeg","isPro":false,"fullname":"Yijia Xiao","user":"Yijia-Xiao","type":"user","name":"Yijia-Xiao"},"name":"Yijia Xiao","status":"claimed_verified","statusLastChangedAt":"2026-08-08T16:45:04.607Z","hidden":false},{"_id":"6847673b3ec10bdd8ab4dd5f","name":"Edward Sun","hidden":false},{"_id":"6847673b3ec10bdd8ab4dd60","name":"Di Luo","hidden":false},{"_id":"6847673b3ec10bdd8ab4dd61","name":"Wei Wang","hidden":false}],"publishedAt":"2024-12-28T12:54:06.000Z","title":"TradingAgents: Multi-Agents LLM Financial Trading Framework","summary":"Significant progress has been made in automated problem-solving using\nsocieties of agents powered by large language models (LLMs). In finance,\nefforts have largely focused on single-agent systems handling specific tasks or\nmulti-agent frameworks independently gathering data. However, the multi-agent\nsystems' potential to replicate real-world trading firms' collaborative\ndynamics remains underexplored. TradingAgents proposes a novel stock trading\nframework inspired by trading firms, featuring LLM-powered agents in\nspecialized roles such as fundamental analysts, sentiment analysts, technical\nanalysts, and traders with varied risk profiles. The framework includes Bull\nand Bear researcher agents assessing market conditions, a risk management team\nmonitoring exposure, and traders synthesizing insights from debates and\nhistorical data to make informed decisions. By simulating a dynamic,\ncollaborative trading environment, this framework aims to improve trading\nperformance. Detailed architecture and extensive experiments reveal its\nsuperiority over baseline models, with notable improvements in cumulative\nreturns, Sharpe ratio, and maximum drawdown, highlighting the potential of\nmulti-agent LLM frameworks in financial trading. TradingAgents is available at\nhttps://github.com/TauricResearch/TradingAgents.","upvotes":123,"discussionId":"6847673c3ec10bdd8ab4dd62","githubRepo":"https://github.com/tauricresearch/tradingagents","githubRepoAddedBy":"auto","ai_summary":"A multi-agent framework using large language models for stock trading simulates real-world trading firms, improving performance metrics like cumulative returns and Sharpe ratio.","ai_keywords":["large language models","LLM","multi-agent systems","fundamental analysts","sentiment analysts","technical analysts","traders","risk management","cumulative returns","Sharpe ratio","maximum drawdown"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":98556},"publishedAt":"2024-12-28T07:54:06.000Z","title":"TradingAgents: Multi-Agents LLM Financial Trading Framework","summary":"Significant progress has been made in automated problem-solving using\nsocieties of agents powered by large language models (LLMs). In finance,\nefforts have largely focused on single-agent systems handling specific tasks or\nmulti-agent frameworks independently gathering data. However, the multi-agent\nsystems' potential to replicate real-world trading firms' collaborative\ndynamics remains underexplored. TradingAgents proposes a novel stock trading\nframework inspired by trading firms, featuring LLM-powered agents in\nspecialized roles such as fundamental analysts, sentiment analysts, technical\nanalysts, and traders with varied risk profiles. The framework includes Bull\nand Bear researcher agents assessing market conditions, a risk management team\nmonitoring exposure, and traders synthesizing insights from debates and\nhistorical data to make informed decisions. By simulating a dynamic,\ncollaborative trading environment, this framework aims to improve trading\nperformance. Detailed architecture and extensive experiments reveal its\nsuperiority over baseline models, with notable improvements in cumulative\nreturns, Sharpe ratio, and maximum drawdown, highlighting the potential of\nmulti-agent LLM frameworks in financial trading. TradingAgents is available at\nhttps://github.com/TauricResearch/TradingAgents.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2412.20138.png","numComments":4,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2310.10688","authors":[{"_id":"663473358043a4686ddb4edb","name":"Abhimanyu Das","hidden":false},{"_id":"663473358043a4686ddb4edc","name":"Weihao Kong","hidden":false},{"_id":"663473358043a4686ddb4edd","name":"Rajat Sen","hidden":false},{"_id":"663473358043a4686ddb4ede","name":"Yichen Zhou","hidden":false}],"publishedAt":"2023-10-14T17:01:37.000Z","title":"A decoder-only foundation model for time-series forecasting","summary":"Motivated by recent advances in large language models for Natural Language\nProcessing (NLP), we design a time-series foundation model for forecasting\nwhose out-of-the-box zero-shot performance on a variety of public datasets\ncomes close to the accuracy of state-of-the-art supervised forecasting models\nfor each individual dataset. Our model is based on pretraining a\npatched-decoder style attention model on a large time-series corpus, and can\nwork well across different forecasting history lengths, prediction lengths and\ntemporal granularities.","upvotes":37,"discussionId":"663473378043a4686ddb4f15","githubRepo":"https://github.com/google-research/timesfm","githubRepoAddedBy":"auto","ai_summary":"A large language model adapted for time-series forecasting achieves near-optimal zero-shot performance on diverse datasets across different time scales and granularities.","ai_keywords":["patched-decoder","attention model","time-series corpus","forecasting","zero-shot performance"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":27839},"publishedAt":"2023-10-14T13:01:37.000Z","title":"A decoder-only foundation model for time-series forecasting","summary":"Motivated by recent advances in large language models for Natural Language\nProcessing (NLP), we design a time-series foundation model for forecasting\nwhose out-of-the-box zero-shot performance on a variety of public datasets\ncomes close to the accuracy of state-of-the-art supervised forecasting models\nfor each individual dataset. Our model is based on pretraining a\npatched-decoder style attention model on a large time-series corpus, and can\nwork well across different forecasting history lengths, prediction lengths and\ntemporal granularities.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2310.10688.png","numComments":1,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2508.02739","authors":[{"_id":"68940073741a16f544fbce2f","name":"Yu Shi","hidden":false},{"_id":"68940073741a16f544fbce30","name":"Zongliang Fu","hidden":false},{"_id":"68940073741a16f544fbce31","name":"Shuo Chen","hidden":false},{"_id":"68940073741a16f544fbce32","name":"Bohan Zhao","hidden":false},{"_id":"68940073741a16f544fbce33","name":"Wei Xu","hidden":false},{"_id":"68940073741a16f544fbce34","name":"Changshui Zhang","hidden":false},{"_id":"68940073741a16f544fbce35","name":"Jian Li","hidden":false}],"publishedAt":"2025-08-02T13:15:59.000Z","title":"Kronos: A Foundation Model for the Language of Financial Markets","summary":"The success of large-scale pre-training paradigm, exemplified by Large\nLanguage Models (LLMs), has inspired the development of Time Series Foundation\nModels (TSFMs). However, their application to financial candlestick (K-line)\ndata remains limited, often underperforming non-pre-trained architectures.\nMoreover, existing TSFMs often overlook crucial downstream tasks such as\nvolatility prediction and synthetic data generation. To address these\nlimitations, we propose Kronos, a unified, scalable pre-training framework\ntailored to financial K-line modeling. Kronos introduces a specialized\ntokenizer that discretizes continuous market information into token sequences,\npreserving both price dynamics and trade activity patterns. We pre-train Kronos\nusing an autoregressive objective on a massive, multi-market corpus of over 12\nbillion K-line records from 45 global exchanges, enabling it to learn nuanced\ntemporal and cross-asset representations. Kronos excels in a zero-shot setting\nacross a diverse set of financial tasks. On benchmark datasets, Kronos boosts\nprice series forecasting RankIC by 93% over the leading TSFM and 87% over the\nbest non-pre-trained baseline. It also achieves a 9% lower MAE in volatility\nforecasting and a 22% improvement in generative fidelity for synthetic K-line\nsequences. These results establish Kronos as a robust, versatile foundation\nmodel for end-to-end financial time series analysis. Our pre-trained model is\npublicly available at https://github.com/shiyu-coder/Kronos.","upvotes":54,"discussionId":"68940073741a16f544fbce36","githubRepo":"https://github.com/shiyu-coder/Kronos","githubRepoAddedBy":"auto","ai_summary":"Kronos, a specialized pre-training framework for financial K-line data, outperforms existing models in forecasting and synthetic data generation through a unique tokenizer and autoregressive pre-training on a large dataset.","ai_keywords":["Large Language Models","Time Series Foundation Models","financial candlestick","volatility prediction","synthetic data generation","autoregressive objective","token sequences","price dynamics","trade activity patterns","zero-shot setting","price series forecasting","RankIC","MAE","generative fidelity"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":37401},"publishedAt":"2025-08-02T09:15:59.000Z","title":"Kronos: A Foundation Model for the Language of Financial Markets","summary":"The success of large-scale pre-training paradigm, exemplified by Large\nLanguage Models (LLMs), has inspired the development of Time Series Foundation\nModels (TSFMs). However, their application to financial candlestick (K-line)\ndata remains limited, often underperforming non-pre-trained architectures.\nMoreover, existing TSFMs often overlook crucial downstream tasks such as\nvolatility prediction and synthetic data generation. To address these\nlimitations, we propose Kronos, a unified, scalable pre-training framework\ntailored to financial K-line modeling. Kronos introduces a specialized\ntokenizer that discretizes continuous market information into token sequences,\npreserving both price dynamics and trade activity patterns. We pre-train Kronos\nusing an autoregressive objective on a massive, multi-market corpus of over 12\nbillion K-line records from 45 global exchanges, enabling it to learn nuanced\ntemporal and cross-asset representations. Kronos excels in a zero-shot setting\nacross a diverse set of financial tasks. On benchmark datasets, Kronos boosts\nprice series forecasting RankIC by 93% over the leading TSFM and 87% over the\nbest non-pre-trained baseline. It also achieves a 9% lower MAE in volatility\nforecasting and a 22% improvement in generative fidelity for synthetic K-line\nsequences. These results establish Kronos as a robust, versatile foundation\nmodel for end-to-end financial time series analysis. Our pre-trained model is\npublicly available at https://github.com/shiyu-coder/Kronos.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2508.02739.png","numComments":4,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2608.03974","authors":[{"_id":"6a729ba11a375f948521c372","name":"Yicheng Xiao","hidden":false},{"_id":"6a729ba11a375f948521c373","name":"Wenxun Dai","hidden":false},{"_id":"6a729ba11a375f948521c374","name":"Xinran Qin","hidden":false},{"_id":"6a729ba11a375f948521c375","name":"Lin Song","hidden":false},{"_id":"6a729ba11a375f948521c376","name":"Maoquan Zhang","hidden":false},{"_id":"6a729ba11a375f948521c377","name":"Hang Xu","hidden":false},{"_id":"6a729ba11a375f948521c378","name":"Yukang Chen","hidden":false},{"_id":"6a729ba11a375f948521c379","name":"Yitong Li","hidden":false},{"_id":"6a729ba11a375f948521c37a","name":"Guohui Zhang","hidden":false},{"_id":"6a729ba11a375f948521c37b","name":"Yuan Zhang","hidden":false},{"_id":"6a729ba11a375f948521c37c","name":"Xuying Zhang","hidden":false},{"_id":"6a729ba11a375f948521c37d","name":"Tommy Zhang","hidden":false},{"_id":"6a729ba11a375f948521c37e","name":"Jianlong Yuan","hidden":false},{"_id":"6a729ba11a375f948521c37f","name":"Peihao Li","hidden":false},{"_id":"6a729ba11a375f948521c380","name":"Shuai Lu","hidden":false},{"_id":"6a729ba11a375f948521c381","name":"Siming Fu","hidden":false},{"_id":"6a729ba11a375f948521c382","name":"Chuyang Zhao","hidden":false},{"_id":"6a729ba11a375f948521c383","name":"Xin Han","hidden":false},{"_id":"6a729ba11a375f948521c384","name":"Jie Huang","hidden":false},{"_id":"6a729ba11a375f948521c385","name":"Wenbo Li","hidden":false},{"_id":"6a729ba11a375f948521c386","name":"Guoqing Ma","hidden":false},{"_id":"6a729ba11a375f948521c387","user":{"_id":"656db3f53dc1d277e5a64410","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/656db3f53dc1d277e5a64410/9kiY2K3MCRcBDk7MrkTBK.png","isPro":false,"fullname":"Wei Huang","user":"AaronHuangWei","type":"user","name":"AaronHuangWei"},"name":"Wei Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-05T08:45:04.532Z","hidden":false},{"_id":"6a729ba11a375f948521c388","name":"Xiaojuan Qi","hidden":false},{"_id":"6a729ba11a375f948521c389","name":"Haoyang Huang","hidden":false},{"_id":"6a729ba11a375f948521c38a","name":"Nan Duan","hidden":false}],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.","upvotes":95,"discussionId":"6a729ba21a375f948521c38b","githubRepo":"https://github.com/jd-opensource/JoyAI-Video-Edit","githubRepoAddedBy":"user","ai_summary":"JoyAI-Video-Edit is a 16B-parameter autoregressive diffusion framework that enables real-time, open-ended video editing with high source fidelity and long-term temporal consistency on a single GPU.","ai_keywords":["autoregressive diffusion","chunk-wise autoregressive adaptation","Source-Anchored Distribution Matching Distillation","SA-DMD","Long-Horizon Autoregressive Distillation","train-inference mismatch","temporal drift","streaming video editing"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1491,"organization":{"_id":"682a97a154b087448a5504ee","name":"jingdong1","fullname":"jingdong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/682a966182dc4fc3e373e3ed/9xUgXZCij6qGJpLqpR99M.png"}},"publishedAt":"2026-08-03T20:00:00.000Z","title":"JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion","summary":"Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2608.03974.png","numComments":1,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"organization":{"_id":"682a97a154b087448a5504ee","name":"jingdong1","fullname":"jingdong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/682a966182dc4fc3e373e3ed/9xUgXZCij6qGJpLqpR99M.png"},"isAuthorParticipating":false},{"paper":{"id":"2606.23050","authors":[{"_id":"6a3a0c36fdcd3514343bb651","user":{"_id":"66bc1f9c8543e3883f2082db","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66bc1f9c8543e3883f2082db/RnmVK90E4VUIkUKcAG3yz.jpeg","isPro":false,"fullname":"Youyang Yin","user":"HYPERUU","type":"user","name":"HYPERUU"},"name":"Youyang Yin","status":"claimed_verified","statusLastChangedAt":"2026-06-23T13:56:12.255Z","hidden":false},{"_id":"6a3a0c36fdcd3514343bb652","name":"Huanhuan Liu","hidden":false},{"_id":"6a3a0c36fdcd3514343bb653","name":"YY","hidden":false},{"_id":"6a3a0c36fdcd3514343bb654","name":"Qunyi Xie","hidden":false},{"_id":"6a3a0c36fdcd3514343bb655","name":"Chaorun Liu","hidden":false},{"_id":"6a3a0c36fdcd3514343bb656","name":"Shiqi Yang","hidden":false},{"_id":"6a3a0c36fdcd3514343bb657","name":"Shaohua Wang","hidden":false},{"_id":"6a3a0c36fdcd3514343bb658","name":"Zhanlong Liu","hidden":false},{"_id":"6a3a0c36fdcd3514343bb659","name":"Hao Zou","hidden":false},{"_id":"6a3a0c36fdcd3514343bb65a","name":"Jinyue Chen","hidden":false},{"_id":"6a3a0c36fdcd3514343bb65b","name":"Shu Wei","hidden":false},{"_id":"6a3a0c36fdcd3514343bb65c","name":"Jingjing Wu","hidden":false},{"_id":"6a3a0c36fdcd3514343bb65d","name":"Mingxin Huang","hidden":false},{"_id":"6a3a0c36fdcd3514343bb65e","name":"Zhen Wu","hidden":false},{"_id":"6a3a0c36fdcd3514343bb65f","name":"Guibin Wang","hidden":false},{"_id":"6a3a0c36fdcd3514343bb660","name":"Tengyu Du","hidden":false},{"_id":"6a3a0c36fdcd3514343bb661","name":"Lei Jia","hidden":false}],"publishedAt":"2026-06-22T00:00:00.000Z","submittedOnDailyAt":"2026-06-23T00:00:00.000Z","title":"Unlimited OCR Works","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a large language model (LLM) as the decoder allows the model to leverage the prior distribution of language, leading to improved OCR performance. However, the downside is equally evident: as the output sequence lengthens, the accumulated KV cache drives up memory consumption and progressively slows down generation. This stands in stark contrast to humans, who exhibit no such decline in efficiency during long-horizon copying tasks. In this technical report, we propose Unlimited OCR, a model designed to emulate human parsing working memory. Taking DeepSeek OCR as the baseline, we replace all attention layers in the decoder with our proposed Reference Sliding Window Attention (R-SWA), which reduces attention computation costs while maintaining a constant KV cache throughout the entire decoding process. By combining the high compression rate of DeepSeek OCR's encoder with our constant KV cache design, Unlimited OCR can transcribe dozens of pages of documents in a single forward pass under a standard maximum length of 32K. More importantly, R-SWA is a general-purpose parsing attention mechanism - beyond OCR, it is equally applicable to tasks such as ASR, translation, etc. Codes and model weights are publicly available at http://github.com/baidu/Unlimited-OCR.","upvotes":83,"discussionId":"6a3a0c36fdcd3514343bb662","githubRepo":"https://github.com/baidu/Unlimited-OCR","githubRepoAddedBy":"user","ai_summary":"Unlimited OCR introduces Reference Sliding Window Attention to eliminate growing memory consumption during long-sequence OCR tasks, enabling efficient transcription of multiple pages in a single forward pass.","ai_keywords":["end-to-end OCR","large language model","decoder","KV cache","attention layers","Reference Sliding Window Attention","document transcription","sequence length","memory consumption","working memory","parsing attention mechanism","ASR","translation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":23948,"organization":{"_id":"626a6d6b4909b521e1f59ce5","name":"baidu","fullname":"BAIDU","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f187a2cc1c03340ac30498/TYYUxK8xD1AxExFMWqbZD.png"}},"publishedAt":"2026-06-21T20:00:00.000Z","title":"Unlimited OCR Works","summary":"Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a large language model (LLM) as the decoder allows the model to leverage the prior distribution of language, leading to improved OCR performance. However, the downside is equally evident: as the output sequence lengthens, the accumulated KV cache drives up memory consumption and progressively slows down generation. This stands in stark contrast to humans, who exhibit no such decline in efficiency during long-horizon copying tasks. In this technical report, we propose Unlimited OCR, a model designed to emulate human parsing working memory. Taking DeepSeek OCR as the baseline, we replace all attention layers in the decoder with our proposed Reference Sliding Window Attention (R-SWA), which reduces attention computation costs while maintaining a constant KV cache throughout the entire decoding process. By combining the high compression rate of DeepSeek OCR's encoder with our constant KV cache design, Unlimited OCR can transcribe dozens of pages of documents in a single forward pass under a standard maximum length of 32K. More importantly, R-SWA is a general-purpose parsing attention mechanism - beyond OCR, it is equally applicable to tasks such as ASR, translation, etc. Codes and model weights are publicly available at http://github.com/baidu/Unlimited-OCR.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2606.23050.png","numComments":7,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"organization":{"_id":"626a6d6b4909b521e1f59ce5","name":"baidu","fullname":"BAIDU","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f187a2cc1c03340ac30498/TYYUxK8xD1AxExFMWqbZD.png"},"isAuthorParticipating":true},{"paper":{"id":"2407.16741","authors":[{"_id":"66a1b11eca8ee359d60b5b99","user":{"_id":"62c63f31031996c36c848504","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62c63f31031996c36c848504/xZvH6jnozuFLXbtRSU-h1.jpeg","isPro":false,"fullname":"Xingyao Wang","user":"xingyaoww","type":"user","name":"xingyaoww"},"name":"Xingyao Wang","status":"extracted_confirmed","statusLastChangedAt":"2024-07-25T02:51:45.594Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5b9a","user":{"_id":"664aefb990135abe9ba9b8c9","avatarUrl":"/avatars/79658c23555deadcfb097b6ebccbe4a6.svg","isPro":false,"fullname":"Boxuan Li","user":"liboxuanhk","type":"user","name":"liboxuanhk"},"name":"Boxuan Li","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:38:14.907Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5b9b","user":{"_id":"65f89af419e43f235cca48f9","avatarUrl":"/avatars/d1ff695f9a8ee0aeeff324fb7f07a811.svg","isPro":false,"fullname":"Yufan Song","user":"yufansong","type":"user","name":"yufansong"},"name":"Yufan Song","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:51:32.404Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5b9c","user":{"_id":"61ac0a0f57307cf7ae3a35bb","avatarUrl":"/avatars/c3b00cf275b2805861f5b020c9d00639.svg","isPro":false,"fullname":"Frank Xu","user":"frankxu","type":"user","name":"frankxu"},"name":"Frank F. Xu","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:51:55.018Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5b9d","user":{"_id":"63357c608adfa81faf2ac180","avatarUrl":"/avatars/ae0314c644f882251baf59b9134fd36f.svg","isPro":false,"fullname":"Xiangru Tang","user":"RTT1","type":"user","name":"RTT1"},"name":"Xiangru Tang","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:52:01.459Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5b9e","user":{"_id":"64403daae44f30a72323e4ca","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64403daae44f30a72323e4ca/skJ9h0pdNfmE4VbQL8xDR.png","isPro":false,"fullname":"mingchen zhuge","user":"tjpxiaoming","type":"user","name":"tjpxiaoming"},"name":"Mingchen Zhuge","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:52:12.418Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5b9f","user":{"_id":"61568f37272f2d87a99ba884","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61568f37272f2d87a99ba884/lgvkl5f0rEyiQRVU5FE32.png","isPro":false,"fullname":"Jiayi Pan","user":"Jiayi-Pan","type":"user","name":"Jiayi-Pan"},"name":"Jiayi Pan","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:52:19.380Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba0","user":{"_id":"63e7bf7be02ee67e8e53f78d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e7bf7be02ee67e8e53f78d/z9j1PWrurEyIA2JXzpz0i.png","isPro":false,"fullname":"Yueqi Song","user":"yueqis","type":"user","name":"yueqis"},"name":"Yueqi Song","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:52:43.838Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba1","user":{"_id":"664aefb990135abe9ba9b8c9","avatarUrl":"/avatars/79658c23555deadcfb097b6ebccbe4a6.svg","isPro":false,"fullname":"Boxuan Li","user":"liboxuanhk","type":"user","name":"liboxuanhk"},"name":"Bowen Li","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:51:24.121Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba2","user":{"_id":"60cc389a0844fb1605fef405","avatarUrl":"/avatars/ec11f85735e0525439e8821cf6d12e53.svg","isPro":false,"fullname":"Jaskirat Singh","user":"jsingh","type":"user","name":"jsingh"},"name":"Jaskirat Singh","status":"claimed_verified","statusLastChangedAt":"2024-07-26T07:40:52.677Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba3","user":{"_id":"65c19c3172fd754bba256112","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65c19c3172fd754bba256112/cfgYvKfO36PSTFYKQottM.jpeg","isPro":false,"fullname":"Ryan Tran","user":"ryanhoangt","type":"user","name":"ryanhoangt"},"name":"Hoang H. Tran","status":"claimed_verified","statusLastChangedAt":"2024-07-25T08:19:36.107Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba4","name":"Fuqiang Li","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba5","user":{"_id":"62581ef5d47cce537ee25963","avatarUrl":"/avatars/d987c625791dc33c84bf9b51069bf2c8.svg","isPro":false,"fullname":"Ren Ma","user":"renma","type":"user","name":"renma"},"name":"Ren Ma","status":"claimed_verified","statusLastChangedAt":"2025-06-11T08:40:40.377Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba6","name":"Mingzhang Zheng","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba7","user":{"_id":"641fa2e0863b87326f4a7fd5","avatarUrl":"/avatars/7b29c20f942d83c08f63f04a85c15c13.svg","isPro":false,"fullname":"Bill Qian","user":"lilbillbiscuit","type":"user","name":"lilbillbiscuit"},"name":"Bill Qian","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:54:51.696Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5ba8","user":{"_id":"62970df979f193515da13dc0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62970df979f193515da13dc0/A-mgKIcgTXRJ54GCHswTq.jpeg","isPro":false,"fullname":"Daniel","user":"super-dainiu","type":"user","name":"super-dainiu"},"name":"Yanjun Shao","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:55:10.708Z","hidden":true},{"_id":"66a1b11eca8ee359d60b5ba9","user":{"_id":"5f1eb362eec0ad2a071ad6e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5f1eb362eec0ad2a071ad6e2/nDiBXdLrOTw67lJp_y_WA.jpeg","isPro":false,"fullname":"Niklas Muennighoff","user":"Muennighoff","type":"user","name":"Muennighoff"},"name":"Niklas Muennighoff","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:55:16.299Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5baa","name":"Yizhe Zhang","hidden":false},{"_id":"66a1b11eca8ee359d60b5bab","user":{"_id":"61e4c4ca1ab24785ac11ba69","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61e4c4ca1ab24785ac11ba69/1Q1zhhyGSJ9RJG9MzwxVv.jpeg","isPro":false,"fullname":"Binyuan Hui","user":"huybery","type":"user","name":"huybery"},"name":"Binyuan Hui","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:55:45.961Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5bac","user":{"_id":"620760a26e3b7210c2ff1943","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620760a26e3b7210c2ff1943/VC-rKqimF6yxGESNVlPoR.jpeg","isPro":false,"fullname":"Junyang Lin","user":"JustinLin610","type":"user","name":"JustinLin610"},"name":"Junyang Lin","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:56:17.993Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5bad","name":"Robert Brennan","hidden":false},{"_id":"66a1b11eca8ee359d60b5bae","user":{"_id":"660ec5a2509153ca49775a7c","avatarUrl":"/avatars/97570fc245cc8ec7628da9c13bd35b71.svg","isPro":false,"fullname":"Hao Peng","user":"haopeng01","type":"user","name":"haopeng01"},"name":"Hao Peng","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:57:13.319Z","hidden":false},{"_id":"66a1b11eca8ee359d60b5baf","name":"Heng Ji","hidden":false},{"_id":"66a1b11eca8ee359d60b5bb0","user":{"_id":"60de14638bedd2315529d43f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1625166923504-noauth.png","isPro":false,"fullname":"Graham Neubig","user":"gneubig","type":"user","name":"gneubig"},"name":"Graham Neubig","status":"admin_assigned","statusLastChangedAt":"2024-07-25T07:57:35.957Z","hidden":false}],"publishedAt":"2024-07-23T17:50:43.000Z","submittedOnDailyAt":"2024-07-25T00:00:00.000Z","title":"OpenDevin: An Open Platform for AI Software Developers as Generalist\n Agents","submittedOnDailyBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","isPro":false,"fullname":"AK","user":"akhaliq","type":"user","name":"akhaliq"},"summary":"Software is one of the most powerful tools that we humans have at our\ndisposal; it allows a skilled programmer to interact with the world in complex\nand profound ways. At the same time, thanks to improvements in large language\nmodels (LLMs), there has also been a rapid development in AI agents that\ninteract with and affect change in their surrounding environments. In this\npaper, we introduce OpenDevin, a platform for the development of powerful and\nflexible AI agents that interact with the world in similar ways to those of a\nhuman developer: by writing code, interacting with a command line, and browsing\nthe web. We describe how the platform allows for the implementation of new\nagents, safe interaction with sandboxed environments for code execution,\ncoordination between multiple agents, and incorporation of evaluation\nbenchmarks. Based on our currently incorporated benchmarks, we perform an\nevaluation of agents over 15 challenging tasks, including software engineering\n(e.g., SWE-Bench) and web browsing (e.g., WebArena), among others. Released\nunder the permissive MIT license, OpenDevin is a community project spanning\nacademia and industry with more than 1.3K contributions from over 160\ncontributors and will improve going forward.","upvotes":84,"discussionId":"66a1b11fca8ee359d60b5bd5","githubRepo":"https://github.com/opendevin/opendevin","githubRepoAddedBy":"auto","ai_summary":"OpenDevin is a platform for developing AI agents that interact with the world by writing code, using command lines, and browsing the web, with support for multiple agents and evaluation benchmarks.","ai_keywords":["large language models","OpenDevin","AI agents","code execution","sandboxed environments","evaluation benchmarks","SWE-Bench","WebArena"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":84268},"publishedAt":"2024-07-23T13:50:43.000Z","title":"OpenDevin: An Open Platform for AI Software Developers as Generalist\n Agents","summary":"Software is one of the most powerful tools that we humans have at our\ndisposal; it allows a skilled programmer to interact with the world in complex\nand profound ways. At the same time, thanks to improvements in large language\nmodels (LLMs), there has also been a rapid development in AI agents that\ninteract with and affect change in their surrounding environments. In this\npaper, we introduce OpenDevin, a platform for the development of powerful and\nflexible AI agents that interact with the world in similar ways to those of a\nhuman developer: by writing code, interacting with a command line, and browsing\nthe web. We describe how the platform allows for the implementation of new\nagents, safe interaction with sandboxed environments for code execution,\ncoordination between multiple agents, and incorporation of evaluation\nbenchmarks. Based on our currently incorporated benchmarks, we perform an\nevaluation of agents over 15 challenging tasks, including software engineering\n(e.g., SWE-Bench) and web browsing (e.g., WebArena), among others. Released\nunder the permissive MIT license, OpenDevin is a community project spanning\nacademia and industry with more than 1.3K contributions from over 160\ncontributors and will improve going forward.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2407.16741.png","numComments":7,"upvoted":false,"submittedBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","fullname":"AK","name":"akhaliq","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":9966,"isUserFollowing":false},"isAuthorParticipating":true},{"paper":{"id":"2608.04205","authors":[{"_id":"6a794cfd8e9301703eaa5edf","name":"Xiaomin Li","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee0","name":"Yuexing Hao","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee1","name":"Jianheng Hou","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee2","name":"Jintao Huang","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee3","name":"Qianfeng Wen","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee4","name":"Shirley Huang","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee5","name":"Yifan Liu","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee6","name":"Xiaoyi Liu","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee7","name":"Yilan Fan","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee8","name":"Yijun Wang","hidden":false},{"_id":"6a794cfd8e9301703eaa5ee9","name":"Koutian Wu","hidden":false},{"_id":"6a794cfd8e9301703eaa5eea","name":"Ruoqi Gao","hidden":false},{"_id":"6a794cfd8e9301703eaa5eeb","user":{"_id":"68eee94b5730be02088503f5","avatarUrl":"/avatars/6b44e0b1fcbbce82787ceb02f0f13fbc.svg","isPro":false,"fullname":"Muhammad Ahmed Mohsin","user":"muahmed7338","type":"user","name":"muahmed7338"},"name":"Muhammad Ahmed Mohsin","status":"claimed_verified","statusLastChangedAt":"2026-08-11T00:45:04.318Z","hidden":false},{"_id":"6a794cfd8e9301703eaa5eec","name":"Jing Tang","hidden":false},{"_id":"6a794cfd8e9301703eaa5eed","name":"Brihi Joshi","hidden":false},{"_id":"6a794cfd8e9301703eaa5eee","user":{"_id":"69f649e4c0b462260106e6ac","avatarUrl":"/avatars/a4adb99e59c91303f0a5671d4bf0eb27.svg","isPro":false,"fullname":"Heming Liu","user":"heming03","type":"user","name":"heming03"},"name":"Heming Liu","status":"claimed_verified","statusLastChangedAt":"2026-08-12T16:45:05.144Z","hidden":false},{"_id":"6a794cfd8e9301703eaa5eef","name":"Zheyuan Deng","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef0","name":"Zonglin Di","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef1","name":"Sankalp Jajee","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef2","name":"Jiuyao Lu","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef3","name":"Zhiwei Zhang","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef4","name":"Saksham Kapoor","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef5","name":"Ishan Gupta","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef6","name":"Yunhan Zhao","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef7","name":"Chanwoo Park","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef8","name":"Yucheng Lu","hidden":false},{"_id":"6a794cfd8e9301703eaa5ef9","name":"Bing Hu","hidden":false},{"_id":"6a794cfd8e9301703eaa5efa","user":{"_id":"699dc9bf019084a18a35989d","avatarUrl":"/avatars/4c4587aeca266e69a498420487f6076f.svg","isPro":false,"fullname":"Weihang Xiao","user":"WeihangXiao","type":"user","name":"WeihangXiao"},"name":"Weihang Xiao","status":"claimed_verified","statusLastChangedAt":"2026-08-11T00:45:04.306Z","hidden":false},{"_id":"6a794cfd8e9301703eaa5efb","name":"Aravind Mohan","hidden":false},{"_id":"6a794cfd8e9301703eaa5efc","name":"Hanwen Xing","hidden":false},{"_id":"6a794cfd8e9301703eaa5efd","name":"Runyu Zhang","hidden":false},{"_id":"6a794cfd8e9301703eaa5efe","name":"Mihir Kulshreshtha","hidden":false},{"_id":"6a794cfd8e9301703eaa5eff","name":"Yuanda Xu","hidden":false},{"_id":"6a794cfd8e9301703eaa5f00","name":"Qianyu Zhu","hidden":false},{"_id":"6a794cfd8e9301703eaa5f01","name":"Dianzhuo Wang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f02","name":"Yuxin Xiao","hidden":false},{"_id":"6a794cfd8e9301703eaa5f03","name":"Bowen Jiang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f04","name":"Yongye Su","hidden":false},{"_id":"6a794cfd8e9301703eaa5f05","name":"Wenhao Chai","hidden":false},{"_id":"6a794cfd8e9301703eaa5f06","name":"Zuxin Liu","hidden":false},{"_id":"6a794cfd8e9301703eaa5f07","name":"Lawrence Yunliang Chen","hidden":false},{"_id":"6a794cfd8e9301703eaa5f08","name":"Xuandong Zhao","hidden":false},{"_id":"6a794cfd8e9301703eaa5f09","name":"Ethan Ye","hidden":false},{"_id":"6a794cfd8e9301703eaa5f0a","name":"Shivam Patel","hidden":false},{"_id":"6a794cfd8e9301703eaa5f0b","name":"Jason Xie","hidden":false},{"_id":"6a794cfd8e9301703eaa5f0c","name":"Alex Martin Richmond","hidden":false},{"_id":"6a794cfd8e9301703eaa5f0d","name":"Weixiang Ding","hidden":false},{"_id":"6a794cfd8e9301703eaa5f0e","name":"Emre Okcular","hidden":false},{"_id":"6a794cfd8e9301703eaa5f0f","name":"Diya Mathew","hidden":false},{"_id":"6a794cfd8e9301703eaa5f10","name":"Ziheng Wang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f11","name":"Rana M. Shahroz Khan","hidden":false},{"_id":"6a794cfd8e9301703eaa5f12","name":"Zhejian Peng","hidden":false},{"_id":"6a794cfd8e9301703eaa5f13","name":"Fang Wu","hidden":false},{"_id":"6a794cfd8e9301703eaa5f14","name":"Fan Nie","hidden":false},{"_id":"6a794cfd8e9301703eaa5f15","name":"Xinyang Han","hidden":false},{"_id":"6a794cfd8e9301703eaa5f16","name":"Yubin Kim","hidden":false},{"_id":"6a794cfd8e9301703eaa5f17","name":"Jiawei Zhang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f18","name":"Zhenting Qi","hidden":false},{"_id":"6a794cfd8e9301703eaa5f19","user":{"_id":"63898e9a2a897944ea5f1453","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898e9a2a897944ea5f1453/EjtslshgnAn5MjKHWLIxi.png","isPro":false,"fullname":"Huangyuan Su","user":"CHL0e","type":"user","name":"CHL0e"},"name":"Huangyuan Su","status":"claimed_verified","statusLastChangedAt":"2026-08-11T00:45:04.313Z","hidden":false},{"_id":"6a794cfd8e9301703eaa5f1a","name":"Xu Pan","hidden":false},{"_id":"6a794cfd8e9301703eaa5f1b","name":"Abinitha Gourabathina","hidden":false},{"_id":"6a794cfd8e9301703eaa5f1c","name":"Hyewon Jeong","hidden":false},{"_id":"6a794cfd8e9301703eaa5f1d","name":"Hemanth Neelgund Ramesh","hidden":false},{"_id":"6a794cfd8e9301703eaa5f1e","name":"Kumail Alhamoud","hidden":false},{"_id":"6a794cfd8e9301703eaa5f1f","name":"Kimia Hamidieh","hidden":false},{"_id":"6a794cfd8e9301703eaa5f20","name":"Zidi Xiong","hidden":false},{"_id":"6a794cfd8e9301703eaa5f21","name":"Samuel Schmidgall","hidden":false},{"_id":"6a794cfd8e9301703eaa5f22","name":"Pengrui Han","hidden":false},{"_id":"6a794cfd8e9301703eaa5f23","name":"Yepeng Huang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f24","name":"Yongheng Wang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f25","name":"Bowen Yang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f26","name":"Alex Gu","hidden":false},{"_id":"6a794cfd8e9301703eaa5f27","name":"Yuchu Wang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f28","name":"Akshay Paruchuri","hidden":false},{"_id":"6a794cfd8e9301703eaa5f29","name":"Brenna Li","hidden":false},{"_id":"6a794cfd8e9301703eaa5f2a","name":"Hejie Cui","hidden":false},{"_id":"6a794cfd8e9301703eaa5f2b","name":"Jiayuan Ding","hidden":false},{"_id":"6a794cfd8e9301703eaa5f2c","name":"Chaosheng Dong","hidden":false},{"_id":"6a794cfd8e9301703eaa5f2d","name":"Jiahao Wang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f2e","name":"Yixuan He","hidden":false},{"_id":"6a794cfd8e9301703eaa5f2f","name":"Chi Wang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f30","name":"Pamela Bhattacharya","hidden":false},{"_id":"6a794cfd8e9301703eaa5f31","name":"Tianyi Peng","hidden":false},{"_id":"6a794cfd8e9301703eaa5f32","name":"Paul Pu Liang","hidden":false},{"_id":"6a794cfd8e9301703eaa5f33","name":"Mitchell Gordon","hidden":false},{"_id":"6a794cfd8e9301703eaa5f34","name":"Yilun Du","hidden":false},{"_id":"6a794cfd8e9301703eaa5f35","name":"Marinka Zitnik","hidden":false},{"_id":"6a794cfd8e9301703eaa5f36","name":"James Zou","hidden":false},{"_id":"6a794cfd8e9301703eaa5f37","name":"Prasanna Tambe","hidden":false},{"_id":"6a794cfd8e9301703eaa5f38","name":"Philip Torr","hidden":false},{"_id":"6a794cfd8e9301703eaa5f39","name":"Emily Fox","hidden":false},{"_id":"6a794cfd8e9301703eaa5f3a","name":"Asu Ozdaglar","hidden":false},{"_id":"6a794cfd8e9301703eaa5f3b","name":"Dawn Song","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/_OwdX_RTgaUMUDHFkbndb.mp4","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/jjMO0eRkai9bdeHzWIPna.jpeg","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/ReOceIa3T1jCopW1VYrlO.jpeg","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/2_rgmAAHiPSrFJ0OjGeW_.png","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/MflY36GOKlAaxg4njTGe1.mp4"],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"MatrAIx: Simulating the World with 8.3 Billion Persona Agents","submittedOnDailyBy":{"_id":"6a3856cf479386f8ed57e427","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a3856cf479386f8ed57e427/1dp3HN5qWeXqYp4knRk4y.png","isPro":true,"fullname":"MatrAIx","user":"MatrAIx","type":"user","name":"MatrAIx"},"summary":"Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.","upvotes":43,"discussionId":"6a794cfe8e9301703eaa5f3c","projectPage":"https://matraix.ai/","githubRepo":"https://github.com/MatrAIx-ai/MatrAIx-Persona-8B","githubRepoAddedBy":"user","ai_summary":"MatrAIx is a large-scale simulated-user evaluation framework that uses diverse persona records and interactive environments to test AI systems across many domains.","ai_keywords":["Persona 8B","dependency graph","coreset","simulated-user evaluation","MatrAIx Playground","persona agents","LLMs"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1152,"organization":{"_id":"6a3875502b67d64201e7b67c","name":"MatrAIx2026","fullname":"MatrAIx","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a3856cf479386f8ed57e427/Z6H0B0aut-U6sv8itS_AW.png"}},"publishedAt":"2026-08-03T20:00:00.000Z","title":"MatrAIx: Simulating the World with 8.3 Billion Persona Agents","summary":"Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.","mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/_OwdX_RTgaUMUDHFkbndb.mp4","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/jjMO0eRkai9bdeHzWIPna.jpeg","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/ReOceIa3T1jCopW1VYrlO.jpeg","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/2_rgmAAHiPSrFJ0OjGeW_.png","https://cdn-uploads.huggingface.co/production/uploads/6a3856cf479386f8ed57e427/MflY36GOKlAaxg4njTGe1.mp4"],"thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2608.04205.png","numComments":3,"upvoted":false,"submittedBy":{"_id":"6a3856cf479386f8ed57e427","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a3856cf479386f8ed57e427/1dp3HN5qWeXqYp4knRk4y.png","fullname":"MatrAIx","name":"MatrAIx","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"organization":{"_id":"6a3875502b67d64201e7b67c","name":"MatrAIx2026","fullname":"MatrAIx","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a3856cf479386f8ed57e427/Z6H0B0aut-U6sv8itS_AW.png"},"isAuthorParticipating":false},{"paper":{"id":"2309.06180","authors":[{"_id":"650114d996ec8a5218886621","user":{"_id":"6458733caa426bae5e079b7b","avatarUrl":"/avatars/990d0c21e28f973290ce429f62156031.svg","isPro":false,"fullname":"Woosuk Kwon","user":"wskwon","type":"user","name":"wskwon"},"name":"Woosuk Kwon","status":"admin_assigned","statusLastChangedAt":"2023-09-13T07:28:58.381Z","hidden":false},{"_id":"650114d996ec8a5218886622","user":{"_id":"649dd1fae16abbd8fd7a3bdf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/2t1p0Ku7qmNcSapRdzxv9.jpeg","isPro":false,"fullname":"Zhuohan Li","user":"zhuohan123","type":"user","name":"zhuohan123"},"name":"Zhuohan Li","status":"admin_assigned","statusLastChangedAt":"2023-09-13T07:29:05.538Z","hidden":false},{"_id":"650114d996ec8a5218886623","name":"Siyuan Zhuang","hidden":false},{"_id":"650114d996ec8a5218886624","user":{"_id":"6455b969654d8bccae50736c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6455b969654d8bccae50736c/WsK0_sbQSVeosi-fh-O_U.jpeg","isPro":false,"fullname":"Ying Sheng","user":"ying1123","type":"user","name":"ying1123"},"name":"Ying Sheng","status":"admin_assigned","statusLastChangedAt":"2023-09-13T07:29:40.177Z","hidden":false},{"_id":"650114d996ec8a5218886625","name":"Lianmin Zheng","hidden":false},{"_id":"650114d996ec8a5218886626","user":{"_id":"64763d68288784ce9aba1cc7","avatarUrl":"/avatars/84033af4466f5469c6264e7bff9f6819.svg","isPro":false,"fullname":"Cody Yu","user":"comaniac","type":"user","name":"comaniac"},"name":"Cody Hao Yu","status":"claimed_verified","statusLastChangedAt":"2023-09-14T07:31:46.459Z","hidden":false},{"_id":"650114d996ec8a5218886627","user":{"_id":"645d2e8401f4eaab2a0878ce","avatarUrl":"/avatars/1273c5fb607b4b622a746a42692fa632.svg","isPro":false,"fullname":"Joseph E. Gonzalez","user":"ProfJoeyG","type":"user","name":"ProfJoeyG"},"name":"Joseph E. Gonzalez","status":"admin_assigned","statusLastChangedAt":"2023-09-13T07:30:30.707Z","hidden":false},{"_id":"650114d996ec8a5218886628","user":{"_id":"62d363143eebd640a4fa41fa","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d363143eebd640a4fa41fa/pvPwXlJ5OOb-UIfmffv4E.jpeg","isPro":false,"fullname":"Hao Zhang","user":"zhisbug","type":"user","name":"zhisbug"},"name":"Hao Zhang","status":"claimed_verified","statusLastChangedAt":"2025-01-01T20:15:04.984Z","hidden":false},{"_id":"650114d996ec8a5218886629","name":"Ion Stoica","hidden":false}],"publishedAt":"2023-09-12T12:50:04.000Z","submittedOnDailyAt":"2023-09-13T00:00:00.000Z","title":"Efficient Memory Management for Large Language Model Serving with\n PagedAttention","submittedOnDailyBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","isPro":false,"fullname":"AK","user":"akhaliq","type":"user","name":"akhaliq"},"summary":"High throughput serving of large language models (LLMs) requires batching\nsufficiently many requests at a time. However, existing systems struggle\nbecause the key-value cache (KV cache) memory for each request is huge and\ngrows and shrinks dynamically. When managed inefficiently, this memory can be\nsignificantly wasted by fragmentation and redundant duplication, limiting the\nbatch size. To address this problem, we propose PagedAttention, an attention\nalgorithm inspired by the classical virtual memory and paging techniques in\noperating systems. On top of it, we build vLLM, an LLM serving system that\nachieves (1) near-zero waste in KV cache memory and (2) flexible sharing of KV\ncache within and across requests to further reduce memory usage. Our\nevaluations show that vLLM improves the throughput of popular LLMs by\n2-4times with the same level of latency compared to the state-of-the-art\nsystems, such as FasterTransformer and Orca. The improvement is more pronounced\nwith longer sequences, larger models, and more complex decoding algorithms.\nvLLM's source code is publicly available at\nhttps://github.com/vllm-project/vllm","upvotes":66,"discussionId":"650114d996ec8a5218886638","githubRepo":"https://github.com/vllm-project/vllm","githubRepoAddedBy":"auto","ai_summary":"PagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache.","ai_keywords":["PagedAttention","key-value cache","KV cache","vLLM","virtual memory","paging techniques","attention algorithm","memory management","throughput improvement","FasterTransformer","Orca"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":86094},"publishedAt":"2023-09-12T08:50:04.000Z","title":"Efficient Memory Management for Large Language Model Serving with\n PagedAttention","summary":"High throughput serving of large language models (LLMs) requires batching\nsufficiently many requests at a time. However, existing systems struggle\nbecause the key-value cache (KV cache) memory for each request is huge and\ngrows and shrinks dynamically. When managed inefficiently, this memory can be\nsignificantly wasted by fragmentation and redundant duplication, limiting the\nbatch size. To address this problem, we propose PagedAttention, an attention\nalgorithm inspired by the classical virtual memory and paging techniques in\noperating systems. On top of it, we build vLLM, an LLM serving system that\nachieves (1) near-zero waste in KV cache memory and (2) flexible sharing of KV\ncache within and across requests to further reduce memory usage. Our\nevaluations show that vLLM improves the throughput of popular LLMs by\n2-4times with the same level of latency compared to the state-of-the-art\nsystems, such as FasterTransformer and Orca. The improvement is more pronounced\nwith longer sequences, larger models, and more complex decoding algorithms.\nvLLM's source code is publicly available at\nhttps://github.com/vllm-project/vllm","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2309.06180.png","numComments":1,"upvoted":false,"submittedBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","fullname":"AK","name":"akhaliq","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":9966,"isUserFollowing":false},"isAuthorParticipating":false},{"paper":{"id":"2502.05512","authors":[{"_id":"67af222945a2cebaa6dfc66d","name":"Wei Deng","hidden":false},{"_id":"67af222945a2cebaa6dfc66e","name":"Siyi Zhou","hidden":false},{"_id":"67af222945a2cebaa6dfc66f","name":"Jingchen Shu","hidden":false},{"_id":"67af222945a2cebaa6dfc670","name":"Jinchao Wang","hidden":false},{"_id":"67af222945a2cebaa6dfc671","name":"Lu Wang","hidden":false}],"publishedAt":"2025-02-08T10:23:20.000Z","title":"IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot\n Text-To-Speech System","summary":"Recently, large language model (LLM) based text-to-speech (TTS) systems have\ngradually become the mainstream in the industry due to their high naturalness\nand powerful zero-shot voice cloning capabilities.Here, we introduce the\nIndexTTS system, which is mainly based on the XTTS and Tortoise model. We add\nsome novel improvements. Specifically, in Chinese scenarios, we adopt a hybrid\nmodeling method that combines characters and pinyin, making the pronunciations\nof polyphonic characters and long-tail characters controllable. We also\nperformed a comparative analysis of the Vector Quantization (VQ) with\nFinite-Scalar Quantization (FSQ) for codebook utilization of acoustic speech\ntokens. To further enhance the effect and stability of voice cloning, we\nintroduce a conformer-based speech conditional encoder and replace the\nspeechcode decoder with BigVGAN2. Compared with XTTS, it has achieved\nsignificant improvements in naturalness, content consistency, and zero-shot\nvoice cloning. As for the popular TTS systems in the open-source, such as\nFish-Speech, CosyVoice2, FireRedTTS and F5-TTS, IndexTTS has a relatively\nsimple training process, more controllable usage, and faster inference speed.\nMoreover, its performance surpasses that of these systems. Our demos are\navailable at https://index-tts.github.io.","upvotes":8,"discussionId":"67af222a45a2cebaa6dfc691","projectPage":"https://index-tts.github.io","githubRepo":"https://github.com/index-tts/index-tts","githubRepoAddedBy":"auto","ai_summary":"IndexTTS, an enhanced text-to-speech system combining XTTS and Tortoise models, offers improved naturalness, enhanced voice cloning, and controllable usage through hybrid character-pinyin modeling and optimized vector quantization.","ai_keywords":["IndexTTS","XTTS","Tortoise","hybrid modeling","characters","pinyin","Vector Quantization (VQ)","Finite-Scalar Quantization (FSQ)","acoustic speech tokens","conformer-based speech conditional encoder","BigVGAN2","naturalness","content consistency","zero-shot voice cloning","Fish-Speech","CosyVoice2","FireRedTTS","F5-TTS"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":23080},"publishedAt":"2025-02-08T05:23:20.000Z","title":"IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot\n Text-To-Speech System","summary":"Recently, large language model (LLM) based text-to-speech (TTS) systems have\ngradually become the mainstream in the industry due to their high naturalness\nand powerful zero-shot voice cloning capabilities.Here, we introduce the\nIndexTTS system, which is mainly based on the XTTS and Tortoise model. We add\nsome novel improvements. Specifically, in Chinese scenarios, we adopt a hybrid\nmodeling method that combines characters and pinyin, making the pronunciations\nof polyphonic characters and long-tail characters controllable. We also\nperformed a comparative analysis of the Vector Quantization (VQ) with\nFinite-Scalar Quantization (FSQ) for codebook utilization of acoustic speech\ntokens. To further enhance the effect and stability of voice cloning, we\nintroduce a conformer-based speech conditional encoder and replace the\nspeechcode decoder with BigVGAN2. Compared with XTTS, it has achieved\nsignificant improvements in naturalness, content consistency, and zero-shot\nvoice cloning. As for the popular TTS systems in the open-source, such as\nFish-Speech, CosyVoice2, FireRedTTS and F5-TTS, IndexTTS has a relatively\nsimple training process, more controllable usage, and faster inference speed.\nMoreover, its performance surpasses that of these systems. Our demos are\navailable at https://index-tts.github.io.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2502.05512.png","numComments":0,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2303.10894","authors":[{"_id":"6940cd6265f1e24a117800af","name":"Xiaoqi Zhao","hidden":false},{"_id":"6940cd6265f1e24a117800b0","name":"Hongpeng Jia","hidden":false},{"_id":"6940cd6265f1e24a117800b1","user":{"_id":"676d078773416936aad01593","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/676d078773416936aad01593/qO-OP0YOZv5cBitgVYl3T.png","isPro":false,"fullname":"yowey","user":"Yoweey","type":"user","name":"Yoweey"},"name":"Youwei Pang","status":"claimed_verified","statusLastChangedAt":"2026-06-18T12:48:37.266Z","hidden":false},{"_id":"6940cd6265f1e24a117800b2","name":"Long Lv","hidden":false},{"_id":"6940cd6265f1e24a117800b3","name":"Feng Tian","hidden":false},{"_id":"6940cd6265f1e24a117800b4","name":"Lihe Zhang","hidden":false},{"_id":"6940cd6265f1e24a117800b5","name":"Weibing Sun","hidden":false},{"_id":"6940cd6265f1e24a117800b6","name":"Huchuan Lu","hidden":false}],"publishedAt":"2023-03-20T06:26:49.000Z","title":"M^{2}SNet: Multi-scale in Multi-scale Subtraction Network for Medical Image Segmentation","summary":"Accurate medical image segmentation is critical for early medical diagnosis. Most existing methods are based on U-shape structure and use element-wise addition or concatenation to fuse different level features progressively in decoder. However, both the two operations easily generate plenty of redundant information, which will weaken the complementarity between different level features, resulting in inaccurate localization and blurred edges of lesions. To address this challenge, we propose a general multi-scale in multi-scale subtraction network (M^{2}SNet) to finish diverse segmentation from medical image. Specifically, we first design a basic subtraction unit (SU) to produce the difference features between adjacent levels in encoder. Next, we expand the single-scale SU to the intra-layer multi-scale SU, which can provide the decoder with both pixel-level and structure-level difference information. Then, we pyramidally equip the multi-scale SUs at different levels with varying receptive fields, thereby achieving the inter-layer multi-scale feature aggregation and obtaining rich multi-scale difference information. In addition, we build a training-free network ``LossNet'' to comprehensively supervise the task-aware features from bottom layer to top layer, which drives our multi-scale subtraction network to capture the detailed and structural cues simultaneously. Without bells and whistles, our method performs favorably against most state-of-the-art methods under different evaluation metrics on eleven datasets of four different medical image segmentation tasks of diverse image modalities, including color colonoscopy imaging, ultrasound imaging, computed tomography (CT), and optical coherence tomography (OCT). The source code can be available at https://github.com/Xiaoqi-Zhao-DLUT/MSNet.","upvotes":0,"discussionId":"6940cd6265f1e24a117800b7","githubRepo":"https://github.com/Xiaoqi-Zhao-DLUT/MSNet","githubRepoAddedBy":"auto","ai_summary":"A multi-scale subtraction network (M$^{2}$SNet) enhances medical image segmentation by capturing detailed and structural cues, improving localization and edge sharpness compared to traditional methods.","ai_keywords":["U-shape structure","feature fusion","element-wise addition","concatenation","redundancy","feature complementarity","subtraction unit (SU)","intra-layer multi-scale SU","inter-layer multi-scale feature aggregation","receptive fields","LossNet","task-aware features","medical image segmentation","color colonoscopy imaging","ultrasound imaging","computed tomography (CT)","optical coherence tomography (OCT)"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":1042},"publishedAt":"2023-03-20T02:26:49.000Z","title":"M^{2}SNet: Multi-scale in Multi-scale Subtraction Network for Medical Image Segmentation","summary":"Accurate medical image segmentation is critical for early medical diagnosis. Most existing methods are based on U-shape structure and use element-wise addition or concatenation to fuse different level features progressively in decoder. However, both the two operations easily generate plenty of redundant information, which will weaken the complementarity between different level features, resulting in inaccurate localization and blurred edges of lesions. To address this challenge, we propose a general multi-scale in multi-scale subtraction network (M^{2}SNet) to finish diverse segmentation from medical image. Specifically, we first design a basic subtraction unit (SU) to produce the difference features between adjacent levels in encoder. Next, we expand the single-scale SU to the intra-layer multi-scale SU, which can provide the decoder with both pixel-level and structure-level difference information. Then, we pyramidally equip the multi-scale SUs at different levels with varying receptive fields, thereby achieving the inter-layer multi-scale feature aggregation and obtaining rich multi-scale difference information. In addition, we build a training-free network ``LossNet'' to comprehensively supervise the task-aware features from bottom layer to top layer, which drives our multi-scale subtraction network to capture the detailed and structural cues simultaneously. Without bells and whistles, our method performs favorably against most state-of-the-art methods under different evaluation metrics on eleven datasets of four different medical image segmentation tasks of diverse image modalities, including color colonoscopy imaging, ultrasound imaging, computed tomography (CT), and optical coherence tomography (OCT). The source code can be available at https://github.com/Xiaoqi-Zhao-DLUT/MSNet.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2303.10894.png","numComments":0,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2509.22186","authors":[{"_id":"68d9ebf80177a6054b013a58","user":{"_id":"650abbb71aece923f21d87fc","avatarUrl":"/avatars/f09ff031c278bc42bfd7a563853e142c.svg","isPro":false,"fullname":"Junbo Niu","user":"Niujunbo2002","type":"user","name":"Niujunbo2002"},"name":"Junbo Niu","status":"claimed_verified","statusLastChangedAt":"2025-10-16T10:38:58.373Z","hidden":false},{"_id":"68d9ebf80177a6054b013a59","user":{"_id":"6625ef13605f46d05c1d0031","avatarUrl":"/avatars/22f201dca35e43013cb593884516e96c.svg","isPro":false,"fullname":"Zheng Liu","user":"starriver030515","type":"user","name":"starriver030515"},"name":"Zheng Liu","status":"claimed_verified","statusLastChangedAt":"2025-09-29T11:26:48.737Z","hidden":false},{"_id":"68d9ebf80177a6054b013a5a","user":{"_id":"658b7d80de33432e4cd795d5","avatarUrl":"/avatars/a83a185edd78157e459b4c0fa6aabfde.svg","isPro":false,"fullname":"gzc","user":"Chokoyo","type":"user","name":"Chokoyo"},"name":"Zhuangcheng Gu","status":"claimed_verified","statusLastChangedAt":"2025-10-01T10:25:10.432Z","hidden":false},{"_id":"68d9ebf80177a6054b013a5b","user":{"_id":"63ae9ff5557befe297a76f90","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1672388558183-noauth.jpeg","isPro":false,"fullname":"Bin Wang","user":"wanderkid","type":"user","name":"wanderkid"},"name":"Bin Wang","status":"claimed_verified","statusLastChangedAt":"2025-09-29T05:58:46.158Z","hidden":false},{"_id":"68d9ebf80177a6054b013a5c","user":{"_id":"66753c556f2ac48ee625d7d1","avatarUrl":"/avatars/8f7c252675fd8a096794d12971903722.svg","isPro":false,"fullname":"Linke Ouyang","user":"ouyanglinke","type":"user","name":"ouyanglinke"},"name":"Linke Ouyang","status":"claimed_verified","statusLastChangedAt":"2025-09-29T11:26:51.347Z","hidden":false},{"_id":"68d9ebf80177a6054b013a5d","name":"Zhiyuan Zhao","hidden":false},{"_id":"68d9ebf80177a6054b013a5e","name":"Tao Chu","hidden":false},{"_id":"68d9ebf80177a6054b013a5f","user":{"_id":"65b8c55130839a0db8cdc496","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b8c55130839a0db8cdc496/REpRsa1ts84yjk1GyyfnT.png","isPro":false,"fullname":"Tianyao He","user":"hotelll","type":"user","name":"hotelll"},"name":"Tianyao He","status":"claimed_verified","statusLastChangedAt":"2025-09-29T11:27:08.370Z","hidden":false},{"_id":"68d9ebf80177a6054b013a60","name":"Fan Wu","hidden":false},{"_id":"68d9ebf80177a6054b013a61","name":"Qintong Zhang","hidden":false},{"_id":"68d9ebf80177a6054b013a62","name":"Zhenjiang Jin","hidden":false},{"_id":"68d9ebf80177a6054b013a63","name":"Guang Liang","hidden":false},{"_id":"68d9ebf80177a6054b013a64","name":"Rui Zhang","hidden":false},{"_id":"68d9ebf80177a6054b013a65","name":"Wenzheng Zhang","hidden":false},{"_id":"68d9ebf80177a6054b013a66","name":"Yuan Qu","hidden":false},{"_id":"68d9ebf80177a6054b013a67","name":"Zhifei Ren","hidden":false},{"_id":"68d9ebf80177a6054b013a68","user":{"_id":"68231f20feab9d28e37f17e1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68231f20feab9d28e37f17e1/hikWkWYg1zmKfU3kLSKeu.jpeg","isPro":false,"fullname":"Yuefeng Sun","user":"SunYuefeng","type":"user","name":"SunYuefeng"},"name":"Yuefeng Sun","status":"claimed_verified","statusLastChangedAt":"2025-09-29T11:27:04.835Z","hidden":false},{"_id":"68d9ebf80177a6054b013a69","name":"Yuanhong Zheng","hidden":false},{"_id":"68d9ebf80177a6054b013a6a","name":"Dongsheng Ma","hidden":false},{"_id":"68d9ebf80177a6054b013a6b","name":"Zirui Tang","hidden":false},{"_id":"68d9ebf80177a6054b013a6c","name":"Boyu Niu","hidden":false},{"_id":"68d9ebf80177a6054b013a6d","name":"Ziyang Miao","hidden":false},{"_id":"68d9ebf80177a6054b013a6e","name":"Hejun Dong","hidden":false},{"_id":"68d9ebf80177a6054b013a6f","name":"Siyi Qian","hidden":false},{"_id":"68d9ebf80177a6054b013a70","user":{"_id":"65411de281a8731a1c28293d","avatarUrl":"/avatars/18607dcd65303e1f688fae8015f43314.svg","isPro":false,"fullname":"junyuan","user":"Carkham","type":"user","name":"Carkham"},"name":"Junyuan Zhang","status":"claimed_verified","statusLastChangedAt":"2026-03-30T09:13:11.612Z","hidden":false},{"_id":"68d9ebf80177a6054b013a71","name":"Jingzhou Chen","hidden":false},{"_id":"68d9ebf80177a6054b013a72","name":"Fangdong Wang","hidden":false},{"_id":"68d9ebf80177a6054b013a73","user":{"_id":"668f77a3b3991ac0c308a441","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/668f77a3b3991ac0c308a441/w5-9UAxi32HnrTjW9_fGC.jpeg","isPro":false,"fullname":"xiaomeng zhao","user":"myhloli","type":"user","name":"myhloli"},"name":"Xiaomeng Zhao","status":"claimed_verified","statusLastChangedAt":"2025-09-30T06:36:34.780Z","hidden":false},{"_id":"68d9ebf80177a6054b013a74","name":"Liqun Wei","hidden":false},{"_id":"68d9ebf80177a6054b013a75","name":"Wei Li","hidden":false},{"_id":"68d9ebf80177a6054b013a76","name":"Shasha Wang","hidden":false},{"_id":"68d9ebf80177a6054b013a77","name":"Ruiliang Xu","hidden":false},{"_id":"68d9ebf80177a6054b013a78","name":"Yuanyuan Cao","hidden":false},{"_id":"68d9ebf80177a6054b013a79","name":"Lu Chen","hidden":false},{"_id":"68d9ebf80177a6054b013a7a","name":"Qianqian Wu","hidden":false},{"_id":"68d9ebf80177a6054b013a7b","name":"Huaiyu Gu","hidden":false},{"_id":"68d9ebf80177a6054b013a7c","name":"Lindong Lu","hidden":false},{"_id":"68d9ebf80177a6054b013a7d","name":"Keming Wang","hidden":false},{"_id":"68d9ebf80177a6054b013a7e","name":"Dechen Lin","hidden":false},{"_id":"68d9ebf80177a6054b013a7f","name":"Guanlin Shen","hidden":false},{"_id":"68d9ebf80177a6054b013a80","user":{"_id":"64ef522242da8d2a897d62da","avatarUrl":"/avatars/03611010d247da66696ac8976d4d3ed3.svg","isPro":false,"fullname":"xuanhe zhou","user":"zhouxh19","type":"user","name":"zhouxh19"},"name":"Xuanhe Zhou","status":"claimed_verified","statusLastChangedAt":"2026-02-27T16:48:01.770Z","hidden":false},{"_id":"68d9ebf80177a6054b013a81","name":"Linfeng Zhang","hidden":false},{"_id":"68d9ebf80177a6054b013a82","user":{"_id":"63859cf3b2906edaf83af9f0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63859cf3b2906edaf83af9f0/kajwuVzd4pDucSPlwghxo.png","isPro":true,"fullname":"Yuhang Zang","user":"yuhangzang","type":"user","name":"yuhangzang"},"name":"Yuhang Zang","status":"claimed_verified","statusLastChangedAt":"2025-09-29T05:58:52.431Z","hidden":false},{"_id":"68d9ebf80177a6054b013a83","name":"Xiaoyi Dong","hidden":false},{"_id":"68d9ebf80177a6054b013a84","user":{"_id":"64b4eec4faa3181a5eab9c46","avatarUrl":"/avatars/bcc9bf5cbf67546ad2b4c9ec8b96ac96.svg","isPro":true,"fullname":"Jiaqi Wang","user":"myownskyW7","type":"user","name":"myownskyW7"},"name":"Jiaqi Wang","status":"claimed_verified","statusLastChangedAt":"2025-09-29T05:58:49.303Z","hidden":false},{"_id":"68d9ebf80177a6054b013a85","name":"Bo Zhang","hidden":false},{"_id":"68d9ebf80177a6054b013a86","name":"Lei Bai","hidden":false},{"_id":"68d9ebf80177a6054b013a87","user":{"_id":"64c9beb2904317f42de06dd8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64c9beb2904317f42de06dd8/he3rxfyzfwEd1vLuK6_o2.jpeg","isPro":false,"fullname":"Pei Chu","user":"chupei","type":"user","name":"chupei"},"name":"Pei Chu","status":"claimed_verified","statusLastChangedAt":"2025-09-29T05:58:41.995Z","hidden":false},{"_id":"68d9ebf80177a6054b013a88","name":"Weijia Li","hidden":false},{"_id":"68d9ebf80177a6054b013a89","name":"Jiang Wu","hidden":false},{"_id":"68d9ebf80177a6054b013a8a","user":{"_id":"643e60d96db6ba8c5ee177ad","avatarUrl":"/avatars/73ac7740e462ba0b53a2f2480d9f1e3e.svg","isPro":false,"fullname":"Lijun Wu","user":"apeters","type":"user","name":"apeters"},"name":"Lijun Wu","status":"claimed_verified","statusLastChangedAt":"2025-10-08T11:49:34.562Z","hidden":false},{"_id":"68d9ebf80177a6054b013a8b","name":"Zhenxiang Li","hidden":false},{"_id":"68d9ebf80177a6054b013a8c","name":"Guangyu Wang","hidden":false},{"_id":"68d9ebf80177a6054b013a8d","name":"Zhongying Tu","hidden":false},{"_id":"68d9ebf80177a6054b013a8e","name":"Chao Xu","hidden":false},{"_id":"68d9ebf80177a6054b013a8f","name":"Kai Chen","hidden":false},{"_id":"68d9ebf80177a6054b013a90","name":"Yu Qiao","hidden":false},{"_id":"68d9ebf80177a6054b013a91","name":"Bowen Zhou","hidden":false},{"_id":"68d9ebf80177a6054b013a92","name":"Dahua Lin","hidden":false},{"_id":"68d9ebf80177a6054b013a93","name":"Wentao Zhang","hidden":false},{"_id":"68d9ebf80177a6054b013a94","name":"Conghui He","hidden":false}],"publishedAt":"2025-09-26T10:45:48.000Z","submittedOnDailyAt":"2025-09-29T00:00:00.000Z","title":"MinerU2.5: A Decoupled Vision-Language Model for Efficient\n High-Resolution Document Parsing","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"We introduce MinerU2.5, a 1.2B-parameter document parsing vision-language\nmodel that achieves state-of-the-art recognition accuracy while maintaining\nexceptional computational efficiency. Our approach employs a coarse-to-fine,\ntwo-stage parsing strategy that decouples global layout analysis from local\ncontent recognition. In the first stage, the model performs efficient layout\nanalysis on downsampled images to identify structural elements, circumventing\nthe computational overhead of processing high-resolution inputs. In the second\nstage, guided by the global layout, it performs targeted content recognition on\nnative-resolution crops extracted from the original image, preserving\nfine-grained details in dense text, complex formulas, and tables. To support\nthis strategy, we developed a comprehensive data engine that generates diverse,\nlarge-scale training corpora for both pretraining and fine-tuning. Ultimately,\nMinerU2.5 demonstrates strong document parsing ability, achieving\nstate-of-the-art performance on multiple benchmarks, surpassing both\ngeneral-purpose and domain-specific models across various recognition tasks,\nwhile maintaining significantly lower computational overhead.","upvotes":177,"discussionId":"68d9ebf80177a6054b013a95","projectPage":"https://opendatalab.github.io/MinerU/","githubRepo":"https://github.com/opendatalab/MinerU","githubRepoAddedBy":"user","ai_summary":"MinerU2.5, a 1.2B-parameter document parsing vision-language model, achieves state-of-the-art recognition accuracy with computational efficiency through a coarse-to-fine parsing strategy.","ai_keywords":["document parsing","vision-language model","coarse-to-fine","two-stage parsing","layout analysis","content recognition","downsampled images","native-resolution crops","data engine","pretraining","fine-tuning","state-of-the-art performance","computational overhead"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":77811},"publishedAt":"2025-09-26T06:45:48.000Z","title":"MinerU2.5: A Decoupled Vision-Language Model for Efficient\n High-Resolution Document Parsing","summary":"We introduce MinerU2.5, a 1.2B-parameter document parsing vision-language\nmodel that achieves state-of-the-art recognition accuracy while maintaining\nexceptional computational efficiency. Our approach employs a coarse-to-fine,\ntwo-stage parsing strategy that decouples global layout analysis from local\ncontent recognition. In the first stage, the model performs efficient layout\nanalysis on downsampled images to identify structural elements, circumventing\nthe computational overhead of processing high-resolution inputs. In the second\nstage, guided by the global layout, it performs targeted content recognition on\nnative-resolution crops extracted from the original image, preserving\nfine-grained details in dense text, complex formulas, and tables. To support\nthis strategy, we developed a comprehensive data engine that generates diverse,\nlarge-scale training corpora for both pretraining and fine-tuning. Ultimately,\nMinerU2.5 demonstrates strong document parsing ability, achieving\nstate-of-the-art performance on multiple benchmarks, surpassing both\ngeneral-purpose and domain-specific models across various recognition tasks,\nwhile maintaining significantly lower computational overhead.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2509.22186.png","numComments":2,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"isAuthorParticipating":false},{"paper":{"id":"2504.19413","authors":[{"_id":"681132ccee133b09d9d25a65","user":{"_id":"63d81cb4635499f2164094bf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1675107437601-noauth.png","isPro":false,"fullname":"Prateek Chhikara","user":"prateekchhikara","type":"user","name":"prateekchhikara"},"name":"Prateek Chhikara","status":"claimed_verified","statusLastChangedAt":"2025-04-30T07:56:27.874Z","hidden":false},{"_id":"681132ccee133b09d9d25a66","name":"Dev Khant","hidden":false},{"_id":"681132ccee133b09d9d25a67","name":"Saket Aryan","hidden":false},{"_id":"681132ccee133b09d9d25a68","name":"Taranjeet Singh","hidden":false},{"_id":"681132ccee133b09d9d25a69","user":{"_id":"64b397bf93679cd00945624e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b397bf93679cd00945624e/yQnWK5uve6pMavzn4bR33.jpeg","isPro":false,"fullname":"Deshraj Yadav","user":"deshrajdry","type":"user","name":"deshrajdry"},"name":"Deshraj Yadav","status":"claimed_verified","statusLastChangedAt":"2025-05-04T10:05:24.817Z","hidden":false}],"publishedAt":"2025-04-28T01:46:35.000Z","submittedOnDailyAt":"2025-04-29T00:00:00.000Z","title":"Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory","submittedOnDailyBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","isPro":false,"fullname":"AK","user":"akhaliq","type":"user","name":"akhaliq"},"summary":"Large Language Models (LLMs) have demonstrated remarkable prowess in\ngenerating contextually coherent responses, yet their fixed context windows\npose fundamental challenges for maintaining consistency over prolonged\nmulti-session dialogues. We introduce Mem0, a scalable memory-centric\narchitecture that addresses this issue by dynamically extracting,\nconsolidating, and retrieving salient information from ongoing conversations.\nBuilding on this foundation, we further propose an enhanced variant that\nleverages graph-based memory representations to capture complex relational\nstructures among conversational elements. Through comprehensive evaluations on\nLOCOMO benchmark, we systematically compare our approaches against six baseline\ncategories: (i) established memory-augmented systems, (ii) retrieval-augmented\ngeneration (RAG) with varying chunk sizes and k-values, (iii) a full-context\napproach that processes the entire conversation history, (iv) an open-source\nmemory solution, (v) a proprietary model system, and (vi) a dedicated memory\nmanagement platform. Empirical results show that our methods consistently\noutperform all existing memory systems across four question categories:\nsingle-hop, temporal, multi-hop, and open-domain. Notably, Mem0 achieves 26%\nrelative improvements in the LLM-as-a-Judge metric over OpenAI, while Mem0 with\ngraph memory achieves around 2% higher overall score than the base\nconfiguration. Beyond accuracy gains, we also markedly reduce computational\noverhead compared to full-context method. In particular, Mem0 attains a 91%\nlower p95 latency and saves more than 90% token cost, offering a compelling\nbalance between advanced reasoning capabilities and practical deployment\nconstraints. Our findings highlight critical role of structured, persistent\nmemory mechanisms for long-term conversational coherence, paving the way for\nmore reliable and efficient LLM-driven AI agents.","upvotes":71,"discussionId":"681132cdee133b09d9d25a97","projectPage":"https://mem0.ai/research","githubRepo":"https://github.com/mem0ai/mem0","githubRepoAddedBy":"auto","ai_summary":"Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.","ai_keywords":["Mem0","memory-centric architecture","graph-based memory","salient information","conversational elements","LOCOMO benchmark","memory-augmented systems","retrieval-augmented generation","full-context approach","memory solution","model system","memory management platform","single-hop","temporal","multi-hop","open-domain","LLM-as-a-Judge","reasoning capabilities","deployment constraints","structured memory","persistent memory"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":63401},"publishedAt":"2025-04-27T21:46:35.000Z","title":"Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory","summary":"Large Language Models (LLMs) have demonstrated remarkable prowess in\ngenerating contextually coherent responses, yet their fixed context windows\npose fundamental challenges for maintaining consistency over prolonged\nmulti-session dialogues. We introduce Mem0, a scalable memory-centric\narchitecture that addresses this issue by dynamically extracting,\nconsolidating, and retrieving salient information from ongoing conversations.\nBuilding on this foundation, we further propose an enhanced variant that\nleverages graph-based memory representations to capture complex relational\nstructures among conversational elements. Through comprehensive evaluations on\nLOCOMO benchmark, we systematically compare our approaches against six baseline\ncategories: (i) established memory-augmented systems, (ii) retrieval-augmented\ngeneration (RAG) with varying chunk sizes and k-values, (iii) a full-context\napproach that processes the entire conversation history, (iv) an open-source\nmemory solution, (v) a proprietary model system, and (vi) a dedicated memory\nmanagement platform. Empirical results show that our methods consistently\noutperform all existing memory systems across four question categories:\nsingle-hop, temporal, multi-hop, and open-domain. Notably, Mem0 achieves 26%\nrelative improvements in the LLM-as-a-Judge metric over OpenAI, while Mem0 with\ngraph memory achieves around 2% higher overall score than the base\nconfiguration. Beyond accuracy gains, we also markedly reduce computational\noverhead compared to full-context method. In particular, Mem0 attains a 91%\nlower p95 latency and saves more than 90% token cost, offering a compelling\nbalance between advanced reasoning capabilities and practical deployment\nconstraints. Our findings highlight critical role of structured, persistent\nmemory mechanisms for long-term conversational coherence, paving the way for\nmore reliable and efficient LLM-driven AI agents.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2504.19413.png","numComments":2,"upvoted":false,"submittedBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","fullname":"AK","name":"akhaliq","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":9966,"isUserFollowing":false},"isAuthorParticipating":false},{"paper":{"id":"2606.03264","authors":[{"_id":"6a1fc6bae292c1c78ecb1598","name":"Zelun Zhang","hidden":false},{"_id":"6a1fc6bae292c1c78ecb1599","name":"Hongen Liu","hidden":false},{"_id":"6a1fc6bae292c1c78ecb159a","name":"Suyin Liang","hidden":false},{"_id":"6a1fc6bae292c1c78ecb159b","user":{"_id":"684ad4f6eb7d8ee8f6a92a3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FnzQWz6ZDaLHydcdgkEbI.png","isPro":false,"fullname":"yubo","user":"zhangyubo0722","type":"user","name":"zhangyubo0722"},"name":"Yubo Zhang","status":"claimed_verified","statusLastChangedAt":"2026-06-15T13:21:35.662Z","hidden":false},{"_id":"6a1fc6bae292c1c78ecb159c","name":"Yiqing Xiang","hidden":false},{"_id":"6a1fc6bae292c1c78ecb159d","name":"Jiaxuan Liu","hidden":false},{"_id":"6a1fc6bae292c1c78ecb159e","name":"Ting Sun","hidden":false},{"_id":"6a1fc6bae292c1c78ecb159f","user":{"_id":"67d96e68939c3823ab2e06a5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FJx_Nf_5aD1w1s0RqbyGn.png","isPro":false,"fullname":"Manhui Lin","user":"gggdddfff","type":"user","name":"gggdddfff"},"name":"Manhui Lin","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.152Z","hidden":false},{"_id":"6a1fc6bae292c1c78ecb15a0","name":"Yue Zhang","hidden":false},{"_id":"6a1fc6bae292c1c78ecb15a1","name":"Changda Zhou","hidden":false},{"_id":"6a1fc6bae292c1c78ecb15a2","user":{"_id":"681c1ecd9539bdde5ae1733c","avatarUrl":"/avatars/23322876ba0812bdabc288cc2381a776.svg","isPro":false,"fullname":"Tingquan Gao","user":"Tingquan","type":"user","name":"Tingquan"},"name":"Tingquan Gao","status":"claimed_verified","statusLastChangedAt":"2026-08-05T16:45:04.472Z","hidden":false},{"_id":"6a1fc6bae292c1c78ecb15a3","user":{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","isPro":false,"fullname":"cuicheng","user":"ChengCui","type":"user","name":"ChengCui"},"name":"Cheng Cui","status":"claimed_verified","statusLastChangedAt":"2026-06-03T14:18:57.782Z","hidden":false},{"_id":"6a1fc6bae292c1c78ecb15a4","name":"Yi Liu","hidden":false},{"_id":"6a1fc6bae292c1c78ecb15a5","name":"Dianhai Yu","hidden":false},{"_id":"6a1fc6bae292c1c78ecb15a6","name":"Yanjun Ma","hidden":false}],"publishedAt":"2026-06-02T00:00:00.000Z","submittedOnDailyAt":"2026-06-03T00:00:00.000Z","title":"PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training","submittedOnDailyBy":{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","isPro":false,"fullname":"cuicheng","user":"ChengCui","type":"user","name":"ChengCui"},"summary":"We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining errors concentrate in under-optimized regions where model behavior is unstable, data coverage is sparse, or supervision is unreliable. Rather than expanding the training corpus indiscriminately, PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to these regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score of 96.33% on OmniDocBench v1.6, demonstrates strong competitiveness against top-tier VLMs, and provides a practical post-training recipe for the PaddleOCR-VL series.","upvotes":27,"discussionId":"6a1fc6bbe292c1c78ecb15a7","projectPage":"https://www.paddleocr.com","githubRepo":"https://github.com/PaddlePaddle/PaddleOCR","githubRepoAddedBy":"user","ai_summary":"PaddleOCR-VL-1.6 enhances document parsing performance through targeted data optimization and progressive post-training techniques, achieving state-of-the-art results on OmniDocBench v1.6.","ai_keywords":["document parsing","data optimization","post-training","reinforcement learning","VLMs","OmniDocBench"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":87779,"organization":{"_id":"62067d5d3906f102bc9658bd","name":"PaddlePaddle","fullname":"PaddlePaddle","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1654942635336-5f3ff69679c1ba4c353d0c5a.png"}},"publishedAt":"2026-06-01T20:00:00.000Z","title":"PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training","summary":"We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining errors concentrate in under-optimized regions where model behavior is unstable, data coverage is sparse, or supervision is unreliable. Rather than expanding the training corpus indiscriminately, PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to these regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score of 96.33% on OmniDocBench v1.6, demonstrates strong competitiveness against top-tier VLMs, and provides a practical post-training recipe for the PaddleOCR-VL series.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2606.03264.png","numComments":1,"upvoted":false,"submittedBy":{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","fullname":"cuicheng","name":"ChengCui","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":40,"isUserFollowing":false},"organization":{"_id":"62067d5d3906f102bc9658bd","name":"PaddlePaddle","fullname":"PaddlePaddle","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1654942635336-5f3ff69679c1ba4c353d0c5a.png"},"isAuthorParticipating":true},{"paper":{"id":"2601.03233","authors":[{"_id":"695dc6d9c03d6d81e4399e85","user":{"_id":"6303cc5e0547362a22a51af0","avatarUrl":"/avatars/8f3348f121565bf6c5e1af0e559a43a3.svg","isPro":false,"fullname":"Yoav HaCohen","user":"yoavhacohen","type":"user","name":"yoavhacohen"},"name":"Yoav HaCohen","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:19:29.722Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e86","user":{"_id":"6489c487b9e9258ba065418f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6489c487b9e9258ba065418f/6rzmV3bQ3YxswG6NP2hDW.png","isPro":false,"fullname":"Benny Brazowski","user":"benibraz","type":"user","name":"benibraz"},"name":"Benny Brazowski","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:19:35.981Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e87","user":{"_id":"62dd30a8d43078cd49ac8ad8","avatarUrl":"/avatars/ad599719290637f7817b7508a91c2e2c.svg","isPro":false,"fullname":"Nisan Chiprut","user":"nisan","type":"user","name":"nisan"},"name":"Nisan Chiprut","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:19:41.634Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e88","user":{"_id":"64a7adc087cbd4dc7301fdd6","avatarUrl":"/avatars/b4ec4c3a0409af8ec4a5de05db453034.svg","isPro":false,"fullname":"Yaki Bitterman","user":"jacobitterman","type":"user","name":"jacobitterman"},"name":"Yaki Bitterman","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:19:46.749Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e89","user":{"_id":"65897258509bcae23fa162c9","avatarUrl":"/avatars/29d277a0c425c936e25e82e79caa10a4.svg","isPro":false,"fullname":"Andrew Kvochko","user":"kvochko","type":"user","name":"kvochko"},"name":"Andrew Kvochko","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:19:51.722Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e8a","name":"Avishai Berkowitz","hidden":false},{"_id":"695dc6d9c03d6d81e4399e8b","name":"Daniel Shalem","hidden":false},{"_id":"695dc6d9c03d6d81e4399e8c","user":{"_id":"681af83e2f4aaa88639e703d","avatarUrl":"/avatars/d69b664daad0afb529440c14fdb9bc3a.svg","isPro":false,"fullname":"Daphna Lifschitz","user":"Daphnal","type":"user","name":"Daphnal"},"name":"Daphna Lifschitz","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:20:04.180Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e8d","user":{"_id":"636b97a57631fe5e86fe1fa2","avatarUrl":"/avatars/c568ae26fd4fc2655cd12f15d539db58.svg","isPro":false,"fullname":"Dudu Moshe","user":"dudumoshe","type":"user","name":"dudumoshe"},"name":"Dudu Moshe","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:20:13.512Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e8e","name":"Eitan Porat","hidden":false},{"_id":"695dc6d9c03d6d81e4399e8f","user":{"_id":"677a422979d3c32a5dd87a0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/WUa6E68GpnT2mEMJ41nDd.png","isPro":false,"fullname":"Eitan Richardson","user":"eitanrich","type":"user","name":"eitanrich"},"name":"Eitan Richardson","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:20:22.677Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e90","user":{"_id":"673f6911d83832a6ce15e7bf","avatarUrl":"/avatars/0da6cded3b0e785241a6ba5fdb5d8ceb.svg","isPro":false,"fullname":"Guy Shiran","user":"guysrn","type":"user","name":"guysrn"},"name":"Guy Shiran","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:20:28.250Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e91","user":{"_id":"65744a2fe09de6aa74026d80","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65744a2fe09de6aa74026d80/kCxIKdeBJwAPKmvlm7fDP.jpeg","isPro":false,"fullname":"Itay Chachy","user":"ItayChachy","type":"user","name":"ItayChachy"},"name":"Itay Chachy","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:20:36.781Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e92","name":"Jonathan Chetboun","hidden":false},{"_id":"695dc6d9c03d6d81e4399e93","user":{"_id":"6678365ac411b340b32d6148","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6678365ac411b340b32d6148/7OhHzbu65pa95eYrAbbLW.jpeg","isPro":false,"fullname":"Michael Finkelson","user":"MichaelFinkelson","type":"user","name":"MichaelFinkelson"},"name":"Michael Finkelson","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:20:53.574Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e94","user":{"_id":"6318aa43cb4ca740c4c55651","avatarUrl":"/avatars/24082c776d284393a5a38a99e5c0bab8.svg","isPro":false,"fullname":"michael kupchick","user":"michaellightricks","type":"user","name":"michaellightricks"},"name":"Michael Kupchick","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:20:59.945Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e95","user":{"_id":"673f29b568595672b8d3e90e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673f29b568595672b8d3e90e/4sYADg3mpqMKmJ4fQwaTl.png","isPro":false,"fullname":"Nir Zabari","user":"NirZabariLTX","type":"user","name":"NirZabariLTX"},"name":"Nir Zabari","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:21:07.297Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e96","user":{"_id":"64ae89c043dda9449a1eb1ba","avatarUrl":"/avatars/12cf3de929d38ddd92cc3f3337dc2ed2.svg","isPro":false,"fullname":"Nitzan Guetta","user":"nitzanguetta","type":"user","name":"nitzanguetta"},"name":"Nitzan Guetta","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:21:14.712Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e97","name":"Noa Kotler","hidden":false},{"_id":"695dc6d9c03d6d81e4399e98","user":{"_id":"631f58935ba8c026340b377c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/631f58935ba8c026340b377c/4yoHLdNE99VBb7ji_Mzzj.jpeg","isPro":false,"fullname":"Ofir Bibi","user":"ofirbibi","type":"user","name":"ofirbibi"},"name":"Ofir Bibi","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:21:27.196Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e99","user":{"_id":"674348b46215a2c0878e219b","avatarUrl":"/avatars/8a213e431a1583d1a93377410907c059.svg","isPro":false,"fullname":"Ori Gordon","user":"origordon","type":"user","name":"origordon"},"name":"Ori Gordon","status":"admin_assigned","statusLastChangedAt":"2026-01-07T13:21:34.878Z","hidden":false},{"_id":"695dc6d9c03d6d81e4399e9a","name":"Poriya Panet","hidden":false},{"_id":"695dc6d9c03d6d81e4399e9b","name":"Roi Benita","hidden":false},{"_id":"695dc6d9c03d6d81e4399e9c","name":"Shahar Armon","hidden":false},{"_id":"695dc6d9c03d6d81e4399e9d","name":"Victor Kulikov","hidden":false},{"_id":"695dc6d9c03d6d81e4399e9e","name":"Yaron Inger","hidden":false},{"_id":"695dc6d9c03d6d81e4399e9f","name":"Yonatan Shiftan","hidden":false},{"_id":"695dc6d9c03d6d81e4399ea0","name":"Zeev Melumian","hidden":false},{"_id":"695dc6d9c03d6d81e4399ea1","name":"Zeev Farbman","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/cYoXYuK3pjt85pl5-fvUv.mp4"],"publishedAt":"2026-01-06T18:24:41.000Z","submittedOnDailyAt":"2026-01-07T00:00:00.000Z","title":"LTX-2: Efficient Joint Audio-Visual Foundation Model","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source foundational model capable of generating high-quality, temporally synchronized audiovisual content in a unified manner. LTX-2 consists of an asymmetric dual-stream transformer with a 14B-parameter video stream and a 5B-parameter audio stream, coupled through bidirectional audio-video cross-attention layers with temporal positional embeddings and cross-modality AdaLN for shared timestep conditioning. This architecture enables efficient training and inference of a unified audiovisual model while allocating more capacity for video generation than audio generation. We employ a multilingual text encoder for broader prompt understanding and introduce a modality-aware classifier-free guidance (modality-CFG) mechanism for improved audiovisual alignment and controllability. Beyond generating speech, LTX-2 produces rich, coherent audio tracks that follow the characters, environment, style, and emotion of each scene -- complete with natural background and foley elements. In our evaluations, the model achieves state-of-the-art audiovisual quality and prompt adherence among open-source systems, while delivering results comparable to proprietary models at a fraction of their computational cost and inference time. All model weights and code are publicly released.","upvotes":188,"discussionId":"695dc6d9c03d6d81e4399ea2","projectPage":"https://app.ltx.studio/ltx-2-playground/i2v","githubRepo":"https://github.com/Lightricks/LTX-2","githubRepoAddedBy":"user","ai_summary":"LTX-2 is an open-source audiovisual diffusion model that generates synchronized video and audio content using a dual-stream transformer architecture with cross-modal attention and classifier-free guidance.","ai_keywords":["text-to-video diffusion models","audiovisual content","dual-stream transformer","cross-attention layers","temporal positional embeddings","AdaLN","classifier-free guidance","modality-aware classifier-free guidance","multilingual text encoder","diffusion models"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":9070},"publishedAt":"2026-01-06T13:24:41.000Z","title":"LTX-2: Efficient Joint Audio-Visual Foundation Model","summary":"Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source foundational model capable of generating high-quality, temporally synchronized audiovisual content in a unified manner. LTX-2 consists of an asymmetric dual-stream transformer with a 14B-parameter video stream and a 5B-parameter audio stream, coupled through bidirectional audio-video cross-attention layers with temporal positional embeddings and cross-modality AdaLN for shared timestep conditioning. This architecture enables efficient training and inference of a unified audiovisual model while allocating more capacity for video generation than audio generation. We employ a multilingual text encoder for broader prompt understanding and introduce a modality-aware classifier-free guidance (modality-CFG) mechanism for improved audiovisual alignment and controllability. Beyond generating speech, LTX-2 produces rich, coherent audio tracks that follow the characters, environment, style, and emotion of each scene -- complete with natural background and foley elements. In our evaluations, the model achieves state-of-the-art audiovisual quality and prompt adherence among open-source systems, while delivering results comparable to proprietary models at a fraction of their computational cost and inference time. All model weights and code are publicly released.","mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/cYoXYuK3pjt85pl5-fvUv.mp4"],"thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2601.03233.png","numComments":9,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"isAuthorParticipating":true},{"paper":{"id":"2508.19205","authors":[{"_id":"68ae6a0f364411bea07df70f","user":{"_id":"646c408336505117e22f4b36","avatarUrl":"/avatars/8201da3c4ec51c8e113446f9578d44f6.svg","isPro":false,"fullname":"zhiliang","user":"zzliang","type":"user","name":"zzliang"},"name":"Zhiliang Peng","status":"claimed_verified","statusLastChangedAt":"2025-09-01T07:54:23.745Z","hidden":false},{"_id":"68ae6a0f364411bea07df710","name":"Jianwei Yu","hidden":false},{"_id":"68ae6a0f364411bea07df711","name":"Wenhui Wang","hidden":false},{"_id":"68ae6a0f364411bea07df712","name":"Yaoyao Chang","hidden":false},{"_id":"68ae6a0f364411bea07df713","name":"Yutao Sun","hidden":false},{"_id":"68ae6a0f364411bea07df714","user":{"_id":"5df85abada6d0311fd3d5408","avatarUrl":"/avatars/2331cf703c1b5d3a62e2050b1a6eb108.svg","isPro":false,"fullname":"Li Dong","user":"unilm","type":"user","name":"unilm"},"name":"Li Dong","status":"claimed_verified","statusLastChangedAt":"2025-08-27T07:10:54.323Z","hidden":false},{"_id":"68ae6a0f364411bea07df715","name":"Yi Zhu","hidden":false},{"_id":"68ae6a0f364411bea07df716","name":"Weijiang Xu","hidden":false},{"_id":"68ae6a0f364411bea07df717","name":"Hangbo Bao","hidden":false},{"_id":"68ae6a0f364411bea07df718","name":"Zehua Wang","hidden":false},{"_id":"68ae6a0f364411bea07df719","user":{"_id":"632bd2f72d6a805eeb4bc601","avatarUrl":"/avatars/6e1533e8a599f3068290aa69ac82cab7.svg","isPro":false,"fullname":"HUANG SHAOHAN","user":"buaahsh","type":"user","name":"buaahsh"},"name":"Shaohan Huang","status":"claimed_verified","statusLastChangedAt":"2025-08-28T08:56:32.368Z","hidden":false},{"_id":"68ae6a0f364411bea07df71a","name":"Yan Xia","hidden":false},{"_id":"68ae6a0f364411bea07df71b","name":"Furu Wei","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/5df85abada6d0311fd3d5408/6tNh2DlU1e_nu7eznNjy1.png"],"publishedAt":"2025-08-26T17:09:12.000Z","submittedOnDailyAt":"2025-08-27T00:00:00.000Z","title":"VibeVoice Technical Report","submittedOnDailyBy":{"_id":"5df85abada6d0311fd3d5408","avatarUrl":"/avatars/2331cf703c1b5d3a62e2050b1a6eb108.svg","isPro":false,"fullname":"Li Dong","user":"unilm","type":"user","name":"unilm"},"summary":"This report presents VibeVoice, a novel model designed to synthesize\nlong-form speech with multiple speakers by employing next-token diffusion,\nwhich is a unified method for modeling continuous data by autoregressively\ngenerating latent vectors via diffusion. To enable this, we introduce a novel\ncontinuous speech tokenizer that, when compared to the popular Encodec model,\nimproves data compression by 80 times while maintaining comparable performance.\nThe tokenizer effectively preserves audio fidelity while significantly boosting\ncomputational efficiency for processing long sequences. Thus, VibeVoice can\nsynthesize long-form speech for up to 90 minutes (in a 64K context window\nlength) with a maximum of 4 speakers, capturing the authentic conversational\n``vibe'' and surpassing open-source and proprietary dialogue models.","upvotes":177,"discussionId":"68ae6a0f364411bea07df71c","projectPage":"https://microsoft.github.io/VibeVoice/","githubRepo":"https://github.com/microsoft/VibeVoice","githubRepoAddedBy":"user","ai_summary":"VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.","ai_keywords":["next-token diffusion","continuous speech tokenizer","Encodec","audio fidelity","computational efficiency","long-form speech","multi-speaker synthesis","conversational vibe","dialogue models"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":52766,"organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"}},"publishedAt":"2025-08-26T13:09:12.000Z","title":"VibeVoice Technical Report","summary":"This report presents VibeVoice, a novel model designed to synthesize\nlong-form speech with multiple speakers by employing next-token diffusion,\nwhich is a unified method for modeling continuous data by autoregressively\ngenerating latent vectors via diffusion. To enable this, we introduce a novel\ncontinuous speech tokenizer that, when compared to the popular Encodec model,\nimproves data compression by 80 times while maintaining comparable performance.\nThe tokenizer effectively preserves audio fidelity while significantly boosting\ncomputational efficiency for processing long sequences. Thus, VibeVoice can\nsynthesize long-form speech for up to 90 minutes (in a 64K context window\nlength) with a maximum of 4 speakers, capturing the authentic conversational\n``vibe'' and surpassing open-source and proprietary dialogue models.","mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/5df85abada6d0311fd3d5408/6tNh2DlU1e_nu7eznNjy1.png"],"thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2508.19205.png","numComments":10,"upvoted":false,"submittedBy":{"_id":"5df85abada6d0311fd3d5408","avatarUrl":"/avatars/2331cf703c1b5d3a62e2050b1a6eb108.svg","fullname":"Li Dong","name":"unilm","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":70,"isUserFollowing":false},"organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"},"isAuthorParticipating":true},{"paper":{"id":"1910.03771","authors":[{"_id":"6411c77b6b75ddced388eff8","user":{"_id":"5df7e9e5da6d0311fd3d53f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583857746553-5df7e9e5da6d0311fd3d53f9.jpeg","isPro":true,"fullname":"Thomas Wolf","user":"thomwolf","type":"user","name":"thomwolf"},"name":"Thomas Wolf","status":"claimed_verified","statusLastChangedAt":"2023-03-15T15:10:53.386Z","hidden":false},{"_id":"6411c77b6b75ddced388eff9","user":{"_id":"5e3aec01f55e2b62848a5217","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e3aec01f55e2b62848a5217/PMKS0NNB4MJQlTSFzh918.jpeg","isPro":true,"fullname":"Lysandre","user":"lysandre","type":"user","name":"lysandre"},"name":"Lysandre Debut","status":"claimed_verified","statusLastChangedAt":"2023-03-15T15:10:45.712Z","hidden":false},{"_id":"6411c77b6b75ddced388effa","user":{"_id":"5ecea265968f6028e0559fa5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1619623771844-5ecea265968f6028e0559fa5.jpeg","isPro":true,"fullname":"Victor Sanh","user":"VictorSanh","type":"user","name":"VictorSanh"},"name":"Victor Sanh","status":"claimed_verified","statusLastChangedAt":"2023-03-15T15:11:01.678Z","hidden":false},{"_id":"6411c77b6b75ddced388effb","user":{"_id":"5dd96eb166059660ed1ee413","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/NQtzmrDdbG0H8qkZvRyGk.jpeg","isPro":true,"fullname":"Julien Chaumond","user":"julien-c","type":"user","name":"julien-c"},"name":"Julien Chaumond","status":"claimed_verified","statusLastChangedAt":"2023-03-15T16:15:09.396Z","hidden":false},{"_id":"6411c77b6b75ddced388effc","user":{"_id":"5e67bdd61009063689407479","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583857146757-5e67bdd61009063689407479.jpeg","isPro":true,"fullname":"Clem 🤗","user":"clem","type":"user","name":"clem"},"name":"Clement Delangue","status":"claimed_verified","statusLastChangedAt":"2023-03-15T16:15:02.261Z","hidden":false},{"_id":"6411c77b6b75ddced388effd","user":{"_id":"5e67c0e3100906368940747d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1623071455900-5e67c0e3100906368940747d.png","isPro":false,"fullname":"Anthony Moi","user":"anthony","type":"user","name":"anthony"},"name":"Anthony Moi","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:08:00.526Z","hidden":false},{"_id":"6411c77b6b75ddced388effe","user":{"_id":"5e67de201009063689407481","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1624630689857-5e67de201009063689407481.jpeg","isPro":true,"fullname":"Pierric Cistac","user":"pierric","type":"user","name":"pierric"},"name":"Pierric Cistac","status":"claimed_verified","statusLastChangedAt":"2023-03-15T15:37:08.150Z","hidden":false},{"_id":"6411c77b6b75ddced388efff","name":"Tim Rault","hidden":false},{"_id":"6411c77b6b75ddced388f000","user":{"_id":"5de8d7255c51de1bfc829f99","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5de8d7255c51de1bfc829f99/98fxu2lJMyEsh2j2PtsAs.jpeg","isPro":false,"fullname":"Remi Louf","user":"remi","type":"user","name":"remi"},"name":"Rémi Louf","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:12:01.391Z","hidden":false},{"_id":"6411c77b6b75ddced388f001","user":{"_id":"5e67c47c100906368940747e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583858935715-5e67c47c100906368940747e.jpeg","isPro":true,"fullname":"Morgan Funtowicz","user":"mfuntowicz","type":"user","name":"mfuntowicz"},"name":"Morgan Funtowicz","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:08:17.081Z","hidden":false},{"_id":"6411c77b6b75ddced388f002","user":{"_id":"5e67bdfe100906368940747a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1653638890151-5e67bdfe100906368940747a.jpeg","isPro":false,"fullname":"Joe Davison","user":"joeddav","type":"user","name":"joeddav"},"name":"Joe Davison","status":"claimed_verified","statusLastChangedAt":"2023-03-22T17:33:07.214Z","hidden":false},{"_id":"6411c77b6b75ddced388f003","user":{"_id":"5e67bff7100906368940747c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583859526680-5e67bff7100906368940747c.jpeg","isPro":false,"fullname":"Sam Shleifer","user":"sshleifer","type":"user","name":"sshleifer"},"name":"Sam Shleifer","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:08:43.201Z","hidden":false},{"_id":"6411c77b6b75ddced388f004","user":{"_id":"5dfcb1aada6d0311fd3d5448","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1584435275418-5dfcb1aada6d0311fd3d5448.jpeg","isPro":false,"fullname":"Patrick von Platen","user":"patrickvonplaten","type":"user","name":"patrickvonplaten"},"name":"Patrick von Platen","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:07:44.788Z","hidden":false},{"_id":"6411c77b6b75ddced388f005","user":{"_id":"5e6ae124d4cd9779932a75fd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1584062726416-noauth.jpeg","isPro":false,"fullname":"Clara Ma","user":"clara-ma","type":"user","name":"clara-ma"},"name":"Clara Ma","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:11:27.949Z","hidden":false},{"_id":"6411c77b6b75ddced388f006","user":{"_id":"5ee3a7cd2a3eae3cbdad1305","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1594144055859-5ee3a7cd2a3eae3cbdad1305.jpeg","isPro":false,"fullname":"Yacine Jernite","user":"yjernite","type":"user","name":"yjernite"},"name":"Yacine Jernite","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:09:17.572Z","hidden":false},{"_id":"6411c77b6b75ddced388f007","user":{"_id":"5df8987fda6d0311fd3d540d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1584609257509-5df8987fda6d0311fd3d540d.jpeg","isPro":false,"fullname":"Julien Plu","user":"jplu","type":"user","name":"jplu"},"name":"Julien Plu","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:10:11.675Z","hidden":false},{"_id":"6411c77b6b75ddced388f008","user":{"_id":"5e3c29d8f55e2b62848a5224","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1618318588221-5e3c29d8f55e2b62848a5224.png","isPro":false,"fullname":"Canwen Xu","user":"canwenxu","type":"user","name":"canwenxu"},"name":"Canwen Xu","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:09:54.029Z","hidden":false},{"_id":"6411c77b6b75ddced388f009","user":{"_id":"5e67bed6100906368940747b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583858339980-5e67bed6100906368940747b.jpeg","isPro":false,"fullname":"Teven Le Scao","user":"teven","type":"user","name":"teven"},"name":"Teven Le Scao","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:09:29.235Z","hidden":false},{"_id":"6411c77b6b75ddced388f00a","user":{"_id":"5ef50182b71947201082a4e5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1593126474392-5ef50182b71947201082a4e5.jpeg","isPro":false,"fullname":"Sylvain Gugger","user":"sgugger","type":"user","name":"sgugger"},"name":"Sylvain Gugger","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:10:26.257Z","hidden":false},{"_id":"6411c77b6b75ddced388f00b","user":{"_id":"5e67c530100906368940747f","avatarUrl":"/avatars/49eb2904a913271da68ed6372f3c03bc.svg","isPro":false,"fullname":"Mariama Drame","user":"mariama","type":"user","name":"mariama"},"name":"Mariama Drame","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:10:59.706Z","hidden":false},{"_id":"6411c77b6b75ddced388f00c","user":{"_id":"5e9ecfc04957053f60648a3e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1594214747713-5e9ecfc04957053f60648a3e.png","isPro":true,"fullname":"Quentin Lhoest","user":"lhoestq","type":"user","name":"lhoestq"},"name":"Quentin Lhoest","status":"claimed_verified","statusLastChangedAt":"2023-03-24T10:48:25.319Z","hidden":false},{"_id":"6411c77b6b75ddced388f00d","user":{"_id":"5e694f697127cb58dde42fe1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/J21QWWtFOKqhJTF5qIMlD.jpeg","isPro":false,"fullname":"Sasha Rush","user":"srush","type":"user","name":"srush"},"name":"Alexander M. Rush","status":"admin_assigned","statusLastChangedAt":"2023-03-31T08:09:02.582Z","hidden":false}],"publishedAt":"2019-10-09T03:23:22.000Z","title":"HuggingFace's Transformers: State-of-the-art Natural Language Processing","summary":"Recent progress in natural language processing has been driven by advances in\nboth model architecture and model pretraining. Transformer architectures have\nfacilitated building higher-capacity models and pretraining has made it\npossible to effectively utilize this capacity for a wide variety of tasks.\nTransformers is an open-source library with the goal of opening up\nthese advances to the wider machine learning community. The library consists of\ncarefully engineered state-of-the art Transformer architectures under a unified\nAPI. Backing this library is a curated collection of pretrained models made by\nand available for the community. Transformers is designed to be\nextensible by researchers, simple for practitioners, and fast and robust in\nindustrial deployments. The library is available at\nhttps://github.com/huggingface/transformers.","upvotes":27,"discussionId":"641192313ea54b1aa7e2ec3f","projectPage":"https://huggingface.co","githubRepo":"https://github.com/huggingface/transformers","githubRepoAddedBy":"user","ai_summary":"Transformers library provides state-of-the-art Transformer architectures and pretrained models for natural language processing tasks with a unified API and emphasis on extensibility and robust deployment.","ai_keywords":["Transformer architectures","pretrained models","unified API","extensible","robust deployment"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":164162,"organization":{"_id":"5e67bd5b1009063689407478","name":"huggingface","fullname":"Hugging Face","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583856921041-5dd96eb166059660ed1ee413.png"}},"publishedAt":"2019-10-08T23:23:22.000Z","title":"HuggingFace's Transformers: State-of-the-art Natural Language Processing","summary":"Recent progress in natural language processing has been driven by advances in\nboth model architecture and model pretraining. Transformer architectures have\nfacilitated building higher-capacity models and pretraining has made it\npossible to effectively utilize this capacity for a wide variety of tasks.\nTransformers is an open-source library with the goal of opening up\nthese advances to the wider machine learning community. The library consists of\ncarefully engineered state-of-the art Transformer architectures under a unified\nAPI. Backing this library is a curated collection of pretrained models made by\nand available for the community. Transformers is designed to be\nextensible by researchers, simple for practitioners, and fast and robust in\nindustrial deployments. The library is available at\nhttps://github.com/huggingface/transformers.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/1910.03771.png","numComments":7,"upvoted":false,"organization":{"_id":"5e67bd5b1009063689407478","name":"huggingface","fullname":"Hugging Face","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583856921041-5dd96eb166059660ed1ee413.png"},"isAuthorParticipating":true},{"paper":{"id":"2508.04660","authors":[{"_id":"689c84e2fab6fdd2e52ac9d9","name":"Noah Ziems","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9da","name":"Dilara Soylu","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9db","user":{"_id":"63bf0cbb24e75d1efdfdf8bd","avatarUrl":"/avatars/3bb13e989db31820734cbe6b13be130a.svg","isPro":false,"fullname":"Lakshya A Agrawal","user":"LakshyAAAgrawal","type":"user","name":"LakshyAAAgrawal"},"name":"Lakshya A Agrawal","status":"claimed_verified","statusLastChangedAt":"2026-05-14T11:03:15.784Z","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9dc","name":"Isaac Miller","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9dd","name":"Liheng Lai","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9de","name":"Chen Qian","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9df","name":"Kaiqiang Song","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9e0","name":"Meng Jiang","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9e1","name":"Dan Klein","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9e2","name":"Matei Zaharia","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9e3","name":"Karel D'Oosterlinck","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9e4","name":"Christopher Potts","hidden":false},{"_id":"689c84e2fab6fdd2e52ac9e5","name":"Omar Khattab","hidden":false}],"publishedAt":"2025-08-06T17:28:31.000Z","title":"Multi-module GRPO: Composing Policy Gradients and Prompt Optimization\n for Language Model Programs","summary":"Group Relative Policy Optimization (GRPO) has proven to be an effective tool\nfor post-training language models (LMs). However, AI systems are increasingly\nexpressed as modular programs that mix together multiple LM calls with distinct\nprompt templates and other tools, and it is not clear how best to leverage GRPO\nto improve these systems. We begin to address this challenge by defining\nmmGRPO, a simple multi-module generalization of GRPO that groups LM calls by\nmodule across rollouts and handles variable-length and interrupted\ntrajectories. We find that mmGRPO, composed with automatic prompt optimization,\nimproves accuracy by 11% on average across classification, many-hop search, and\nprivacy-preserving delegation tasks against the post-trained LM, and by 5%\nagainst prompt optimization on its own. We open-source mmGRPO in DSPy as the\ndspy.GRPO optimizer.","upvotes":7,"discussionId":"689c84e3fab6fdd2e52ac9e6","projectPage":"https://dspy.ai","githubRepo":"https://github.com/stanfordnlp/dspy","githubRepoAddedBy":"auto","ai_summary":"mmGRPO, a multi-module extension of GRPO, enhances accuracy in modular AI systems by optimizing LM calls and prompts across various tasks.","ai_keywords":["GRPO","post-training language models","mmGRPO","multi-module","LM calls","prompt templates","automatic prompt optimization","classification","many-hop search","privacy-preserving delegation","dspy.GRPO optimizer"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":37314},"publishedAt":"2025-08-06T13:28:31.000Z","title":"Multi-module GRPO: Composing Policy Gradients and Prompt Optimization\n for Language Model Programs","summary":"Group Relative Policy Optimization (GRPO) has proven to be an effective tool\nfor post-training language models (LMs). However, AI systems are increasingly\nexpressed as modular programs that mix together multiple LM calls with distinct\nprompt templates and other tools, and it is not clear how best to leverage GRPO\nto improve these systems. We begin to address this challenge by defining\nmmGRPO, a simple multi-module generalization of GRPO that groups LM calls by\nmodule across rollouts and handles variable-length and interrupted\ntrajectories. We find that mmGRPO, composed with automatic prompt optimization,\nimproves accuracy by 11% on average across classification, many-hop search, and\nprivacy-preserving delegation tasks against the post-trained LM, and by 5%\nagainst prompt optimization on its own. We open-source mmGRPO in DSPy as the\ndspy.GRPO optimizer.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2508.04660.png","numComments":0,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2503.11576","authors":[{"_id":"67d7d1c38a15934c10b1576d","user":{"_id":"63c64fdd77caf00391006211","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1673940984052-63c64fdd77caf00391006211.jpeg","isPro":true,"fullname":"Ahmed Nassar","user":"asnassar","type":"user","name":"asnassar"},"name":"Ahmed Nassar","status":"claimed_verified","statusLastChangedAt":"2025-03-18T08:09:03.445Z","hidden":false},{"_id":"67d7d1c38a15934c10b1576e","user":{"_id":"65d66b494bbd0d92b641cdbb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65d66b494bbd0d92b641cdbb/6-7dm7B-JxcoS1QlCPdMN.jpeg","isPro":false,"fullname":"Andres Marafioti","user":"andito","type":"user","name":"andito"},"name":"Andres Marafioti","status":"claimed_verified","statusLastChangedAt":"2025-03-17T14:42:18.190Z","hidden":false},{"_id":"67d7d1c38a15934c10b1576f","user":{"_id":"663e1254887b6f5645a0399f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663e1254887b6f5645a0399f/CXtmipcpwK3LMIQNeinVg.jpeg","isPro":false,"fullname":"Matteo Omenetti","user":"MatteoOmenetti","type":"user","name":"MatteoOmenetti"},"name":"Matteo Omenetti","status":"admin_assigned","statusLastChangedAt":"2025-03-17T15:05:38.055Z","hidden":false},{"_id":"67d7d1c38a15934c10b15770","user":{"_id":"665a2d536a8e042a0bf9b7b0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/665a2d536a8e042a0bf9b7b0/wA8QmKPW1BSRWNJyx_st0.jpeg","isPro":false,"fullname":"Maksym Lysak","user":"MaxMnemonic","type":"user","name":"MaxMnemonic"},"name":"Maksym Lysak","status":"claimed_verified","statusLastChangedAt":"2025-03-17T16:32:51.359Z","hidden":false},{"_id":"67d7d1c38a15934c10b15771","user":{"_id":"67938030aa5e507f71eb0e44","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67938030aa5e507f71eb0e44/U_7TxvZFuTeJ-Gtfatbzj.jpeg","isPro":false,"fullname":"Nikos Livathinos","user":"nlivathinos","type":"user","name":"nlivathinos"},"name":"Nikolaos Livathinos","status":"claimed_verified","statusLastChangedAt":"2025-03-17T16:32:49.368Z","hidden":false},{"_id":"67d7d1c38a15934c10b15772","user":{"_id":"67aa428095116c7b8e4204d4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/dT638jwk8CTQ6tGJ0ullR.png","isPro":false,"fullname":"Christoph Auer","user":"ChristophAuer","type":"user","name":"ChristophAuer"},"name":"Christoph Auer","status":"admin_assigned","statusLastChangedAt":"2025-03-17T15:05:12.752Z","hidden":false},{"_id":"67d7d1c38a15934c10b15773","user":{"_id":"64d38f55f8082bf19b7339e0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64d38f55f8082bf19b7339e0/zKbW5sSUY4TulbD20axRg.jpeg","isPro":false,"fullname":"Lucas Morin","user":"lucas-morin","type":"user","name":"lucas-morin"},"name":"Lucas Morin","status":"claimed_verified","statusLastChangedAt":"2025-10-28T15:41:35.917Z","hidden":false},{"_id":"67d7d1c38a15934c10b15774","user":{"_id":"65ce0dbef3bbc55ca4d0528f","avatarUrl":"/avatars/ae19424457dbec49579fe5a4ae7177b0.svg","isPro":false,"fullname":"Rafael Teixeira de Lima","user":"rtdl","type":"user","name":"rtdl"},"name":"Rafael Teixeira de Lima","status":"admin_assigned","statusLastChangedAt":"2025-03-17T15:03:11.221Z","hidden":false},{"_id":"67d7d1c38a15934c10b15775","name":"Yusik Kim","hidden":false},{"_id":"67d7d1c38a15934c10b15776","user":{"_id":"65d4dab8be26e8e084d29b98","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65d4dab8be26e8e084d29b98/67m-1xtaONgP3sJCw5_k1.jpeg","isPro":false,"fullname":"Said Gurbuz","user":"Saidgurbuz","type":"user","name":"Saidgurbuz"},"name":"A. Said Gurbuz","status":"admin_assigned","statusLastChangedAt":"2025-03-17T15:03:01.577Z","hidden":false},{"_id":"67d7d1c38a15934c10b15777","user":{"_id":"63c64dd877caf00391004e20","avatarUrl":"/avatars/6350443da6405ea7beccf47a80d8710e.svg","isPro":false,"fullname":"Michele Dolfi","user":"dolfim-ibm","type":"user","name":"dolfim-ibm"},"name":"Michele Dolfi","status":"admin_assigned","statusLastChangedAt":"2025-03-17T15:02:52.883Z","hidden":false},{"_id":"67d7d1c38a15934c10b15778","user":{"_id":"61ed0ff29539bc0a3bbc89f4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61ed0ff29539bc0a3bbc89f4/iYWK7GParA7Ke5F6q132W.jpeg","isPro":false,"fullname":"Miquel Farré","user":"mfarre","type":"user","name":"mfarre"},"name":"Miquel Farré","status":"admin_assigned","statusLastChangedAt":"2025-03-17T15:05:05.893Z","hidden":false},{"_id":"67d7d1c38a15934c10b15779","user":{"_id":"63c6522d8fddd19431f48bcc","avatarUrl":"/avatars/e2ec9c76656b7520472718595f9dafe1.svg","isPro":false,"fullname":"Peter W. J. Staar","user":"PeterStaar","type":"user","name":"PeterStaar"},"name":"Peter W. J. Staar","status":"admin_assigned","statusLastChangedAt":"2025-03-17T15:01:33.085Z","hidden":false}],"publishedAt":"2025-03-14T16:44:14.000Z","submittedOnDailyAt":"2025-03-17T00:00:00.000Z","title":"SmolDocling: An ultra-compact vision-language model for end-to-end\n multi-modal document conversion","submittedOnDailyBy":{"_id":"65d66b494bbd0d92b641cdbb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65d66b494bbd0d92b641cdbb/6-7dm7B-JxcoS1QlCPdMN.jpeg","isPro":false,"fullname":"Andres Marafioti","user":"andito","type":"user","name":"andito"},"summary":"We introduce SmolDocling, an ultra-compact vision-language model targeting\nend-to-end document conversion. Our model comprehensively processes entire\npages by generating DocTags, a new universal markup format that captures all\npage elements in their full context with location. Unlike existing approaches\nthat rely on large foundational models, or ensemble solutions that rely on\nhandcrafted pipelines of multiple specialized models, SmolDocling offers an\nend-to-end conversion for accurately capturing content, structure and spatial\nlocation of document elements in a 256M parameters vision-language model.\nSmolDocling exhibits robust performance in correctly reproducing document\nfeatures such as code listings, tables, equations, charts, lists, and more\nacross a diverse range of document types including business documents, academic\npapers, technical reports, patents, and forms -- significantly extending beyond\nthe commonly observed focus on scientific papers. Additionally, we contribute\nnovel publicly sourced datasets for charts, tables, equations, and code\nrecognition. Experimental results demonstrate that SmolDocling competes with\nother Vision Language Models that are up to 27 times larger in size, while\nreducing computational requirements substantially. The model is currently\navailable, datasets will be publicly available soon.","upvotes":166,"discussionId":"67d7d1c68a15934c10b15841","projectPage":"https://huggingface.co/ds4sd/SmolDocling-256M-preview","githubRepo":"https://github.com/docling-project/docling","githubRepoAddedBy":"auto","ai_summary":"SmolDocling is a compact vision-language model that performs end-to-end document conversion with robust performance across various document types using 256M parameters and a new markup format.","ai_keywords":["vision-language model","DocTags","end-to-end document conversion","universal markup format","page elements","location context","large foundational models","ensemble solutions","specialized models","code listings","tables","equations","charts","lists","diverse document types","publicly sourced datasets","Vision Language Models","computational requirements"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":64885,"organization":{"_id":"6624c149a8d1362ebc6bc6da","name":"ibm-granite","fullname":"IBM Granite","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/639bcaa2445b133a4e942436/CEW-OjXkRkDNmTxSu8Egh.png"}},"publishedAt":"2025-03-14T12:44:14.000Z","title":"SmolDocling: An ultra-compact vision-language model for end-to-end\n multi-modal document conversion","summary":"We introduce SmolDocling, an ultra-compact vision-language model targeting\nend-to-end document conversion. Our model comprehensively processes entire\npages by generating DocTags, a new universal markup format that captures all\npage elements in their full context with location. Unlike existing approaches\nthat rely on large foundational models, or ensemble solutions that rely on\nhandcrafted pipelines of multiple specialized models, SmolDocling offers an\nend-to-end conversion for accurately capturing content, structure and spatial\nlocation of document elements in a 256M parameters vision-language model.\nSmolDocling exhibits robust performance in correctly reproducing document\nfeatures such as code listings, tables, equations, charts, lists, and more\nacross a diverse range of document types including business documents, academic\npapers, technical reports, patents, and forms -- significantly extending beyond\nthe commonly observed focus on scientific papers. Additionally, we contribute\nnovel publicly sourced datasets for charts, tables, equations, and code\nrecognition. Experimental results demonstrate that SmolDocling competes with\nother Vision Language Models that are up to 27 times larger in size, while\nreducing computational requirements substantially. The model is currently\navailable, datasets will be publicly available soon.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2503.11576.png","numComments":19,"upvoted":false,"submittedBy":{"_id":"65d66b494bbd0d92b641cdbb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65d66b494bbd0d92b641cdbb/6-7dm7B-JxcoS1QlCPdMN.jpeg","fullname":"Andres Marafioti","name":"andito","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":608,"isUserFollowing":false},"organization":{"_id":"6624c149a8d1362ebc6bc6da","name":"ibm-granite","fullname":"IBM Granite","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/639bcaa2445b133a4e942436/CEW-OjXkRkDNmTxSu8Egh.png"},"isAuthorParticipating":true},{"paper":{"id":"2509.06926","authors":[{"_id":"68c02a15629c5c1ac880f540","name":"Rouard Simon","hidden":false},{"_id":"68c02a15629c5c1ac880f541","name":"Orsini Manu","hidden":false},{"_id":"68c02a15629c5c1ac880f542","name":"Roebel Axel","hidden":false},{"_id":"68c02a15629c5c1ac880f543","name":"Zeghidour Neil","hidden":false},{"_id":"68c02a15629c5c1ac880f544","name":"Défossez Alexandre","hidden":false}],"publishedAt":"2025-09-08T17:38:13.000Z","title":"Continuous Audio Language Models","summary":"Audio Language Models (ALM) have emerged as the dominant paradigm for speech\nand music generation by representing audio as sequences of discrete tokens.\nYet, unlike text tokens, which are invertible, audio tokens are extracted from\nlossy codecs with a limited bitrate. As a consequence, increasing audio quality\nrequires generating more tokens, which imposes a trade-off between fidelity and\ncomputational cost. We address this issue by studying Continuous Audio Language\nModels (CALM). These models instantiate a large Transformer backbone that\nproduces a contextual embedding at every timestep. This sequential information\nthen conditions an MLP that generates the next continuous frame of an audio VAE\nthrough consistency modeling. By avoiding lossy compression, CALM achieves\nhigher quality at lower computational cost than their discrete counterpart.\nExperiments on speech and music demonstrate improved efficiency and fidelity\nover state-of-the-art discrete audio language models, facilitating lightweight,\nhigh-quality audio generation. Samples are available at\nhttps://continuous-audio-language-models.github.io","upvotes":12,"discussionId":"68c02a15629c5c1ac880f545","projectPage":"https://huggingface.co/spaces/kyutai/calm-samples","githubRepo":"https://github.com/kyutai-labs/pocket-tts","githubRepoAddedBy":"auto","githubStars":8657},"publishedAt":"2025-09-08T13:38:13.000Z","title":"Continuous Audio Language Models","summary":"Audio Language Models (ALM) have emerged as the dominant paradigm for speech\nand music generation by representing audio as sequences of discrete tokens.\nYet, unlike text tokens, which are invertible, audio tokens are extracted from\nlossy codecs with a limited bitrate. As a consequence, increasing audio quality\nrequires generating more tokens, which imposes a trade-off between fidelity and\ncomputational cost. We address this issue by studying Continuous Audio Language\nModels (CALM). These models instantiate a large Transformer backbone that\nproduces a contextual embedding at every timestep. This sequential information\nthen conditions an MLP that generates the next continuous frame of an audio VAE\nthrough consistency modeling. By avoiding lossy compression, CALM achieves\nhigher quality at lower computational cost than their discrete counterpart.\nExperiments on speech and music demonstrate improved efficiency and fidelity\nover state-of-the-art discrete audio language models, facilitating lightweight,\nhigh-quality audio generation. Samples are available at\nhttps://continuous-audio-language-models.github.io","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2509.06926.png","numComments":0,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2608.01964","authors":[{"_id":"6a715744ec5082b9f872ce05","name":"Ziyu Ma","hidden":false},{"_id":"6a715744ec5082b9f872ce06","user":{"_id":"65003db8bef9b594656f8fa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65003db8bef9b594656f8fa7/L6cvPOAeBRnFnIQwWxYyf.png","isPro":false,"fullname":"Hailang Huang","user":"lerogo","type":"user","name":"lerogo"},"name":"Hailang Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-04T08:45:04.665Z","hidden":false},{"_id":"6a715744ec5082b9f872ce07","name":"Shun Zou","hidden":false},{"_id":"6a715744ec5082b9f872ce08","name":"Yong Wang","hidden":false},{"_id":"6a715744ec5082b9f872ce09","name":"Shidong Yang","hidden":false},{"_id":"6a715744ec5082b9f872ce0a","name":"Yiming Hu","hidden":false},{"_id":"6a715744ec5082b9f872ce0b","name":"Fei Wei","hidden":false},{"_id":"6a715744ec5082b9f872ce0c","name":"XiangXiang Chu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64d1dc5273174cecdffc97d3/QDDq6oSuxJP_g4SwESa7B.mp4"],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks","submittedOnDailyBy":{"_id":"64d1dc5273174cecdffc97d3","avatarUrl":"/avatars/6564e6b68fee9673f75b6366adf39a3b.svg","isPro":false,"fullname":"Wang Yong","user":"seashell11","type":"user","name":"seashell11"},"summary":"Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen~3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench~2.1, and from 2.8% to 8.3% on OSWorld~2.0. It also raises Claude Opus~4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.","upvotes":175,"discussionId":"6a715745ec5082b9f872ce0d","projectPage":"https://lh-harness.pages.dev","githubRepo":"https://github.com/AMAP-ML/LongHorizon-Harness","githubRepoAddedBy":"user","ai_summary":"LongHorizon-Harness improves long-horizon agent performance by explicitly tracking verified task states outside context via a manage-execute-audit loop.","ai_keywords":["large language model agents","long-horizon tasks","task-state management","Manage-Execute-Audit loop","AgentAdapter","WeaveBench","Terminal-Bench","OSWorld"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":800,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"}},"publishedAt":"2026-08-02T20:00:00.000Z","title":"LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks","summary":"Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen~3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench~2.1, and from 2.8% to 8.3% on OSWorld~2.0. It also raises Claude Opus~4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.","mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64d1dc5273174cecdffc97d3/QDDq6oSuxJP_g4SwESa7B.mp4"],"thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2608.01964.png","numComments":3,"upvoted":false,"submittedBy":{"_id":"64d1dc5273174cecdffc97d3","avatarUrl":"/avatars/6564e6b68fee9673f75b6366adf39a3b.svg","fullname":"Wang Yong","name":"seashell11","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"},"isAuthorParticipating":true},{"paper":{"id":"2605.03042","authors":[{"_id":"69fa9fc3cfa95edeb1e79e8a","user":{"_id":"662f43dbe93bb73804fa1606","avatarUrl":"/avatars/1a3c117283598f566fc59b02e0df2f9e.svg","isPro":true,"fullname":"Ruofeng Yang","user":"RuofengYang","type":"user","name":"RuofengYang"},"name":"Ruofeng Yang","status":"claimed_verified","statusLastChangedAt":"2026-05-07T14:22:29.648Z","hidden":false},{"_id":"69fa9fc3cfa95edeb1e79e8b","name":"Yongcan Li","hidden":false},{"_id":"69fa9fc3cfa95edeb1e79e8c","name":"Shuai Li","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/662f43dbe93bb73804fa1606/nFeWKD7V2DRVrzZF_JTDr.png"],"publishedAt":"2026-05-04T00:00:00.000Z","submittedOnDailyAt":"2026-05-06T00:00:00.000Z","title":"ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration","submittedOnDailyBy":{"_id":"662f43dbe93bb73804fa1606","avatarUrl":"/avatars/1a3c117283598f566fc59b02e0df2f9e.svg","isPro":true,"fullname":"Ruofeng Yang","user":"RuofengYang","type":"user","name":"RuofengYang"},"summary":"This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance mechanisms, and early deployment experience. The performance of agent systems built on LLMs depends on both the model weights and the harness around them, which governs what information to store, retrieve, and present to the model. For long-horizon research workflows, the central failure mode is not a visible breakdown but a plausible unsupported success: a long-running agent can produce claims whose evidential support is incomplete, misreported, or silently inherited from the executor's framing. Therefore, we present ARIS as a research harness that coordinates machine-learning research workflows through cross-model adversarial collaboration as a default configuration: an executor model drives forward progress while a reviewer from a different model family is recommended to critique intermediate artifacts and request revisions. ARIS has three architectural layers. The execution layer provides more than 65 reusable Markdown-defined skills, model integrations via MCP, a persistent research wiki for iterative reuse of prior findings, and deterministic figure generation. The orchestration layer coordinates five end-to-end workflows with adjustable effort settings and configurable routing to reviewer models. The assurance layer includes a three-stage process for checking whether experimental claims are supported by evidence: integrity verification, result-to-claim mapping, and claim auditing that cross-checks manuscript statements against the claim ledger and raw evidence, as well as a five-pass scientific-editing pipeline, mathematical-proof checks, and visual inspection of the rendered PDF. A prototype self-improvement loop records research traces and proposes harness improvements that are adopted only after reviewer approval.","upvotes":145,"discussionId":"69fa9fc3cfa95edeb1e79e8d","projectPage":"https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep","githubRepo":"https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep","githubRepoAddedBy":"user","ai_summary":"ARIS is an open-source research harness that uses cross-model adversarial collaboration to ensure reliable long-term research outcomes through coordinated execution, orchestration, and assurance layers.","ai_keywords":["LLMs","agent systems","model weights","research harness","cross-model adversarial collaboration","executor model","reviewer model","Markdown-defined skills","MCP","persistent research wiki","deterministic figure generation","end-to-end workflows","adjustable effort settings","configurable routing","integrity verification","result-to-claim mapping","claim auditing","scientific-editing pipeline","mathematical-proof checks","visual inspection"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":14770,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"publishedAt":"2026-05-03T20:00:00.000Z","title":"ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration","summary":"This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance mechanisms, and early deployment experience. The performance of agent systems built on LLMs depends on both the model weights and the harness around them, which governs what information to store, retrieve, and present to the model. For long-horizon research workflows, the central failure mode is not a visible breakdown but a plausible unsupported success: a long-running agent can produce claims whose evidential support is incomplete, misreported, or silently inherited from the executor's framing. Therefore, we present ARIS as a research harness that coordinates machine-learning research workflows through cross-model adversarial collaboration as a default configuration: an executor model drives forward progress while a reviewer from a different model family is recommended to critique intermediate artifacts and request revisions. ARIS has three architectural layers. The execution layer provides more than 65 reusable Markdown-defined skills, model integrations via MCP, a persistent research wiki for iterative reuse of prior findings, and deterministic figure generation. The orchestration layer coordinates five end-to-end workflows with adjustable effort settings and configurable routing to reviewer models. The assurance layer includes a three-stage process for checking whether experimental claims are supported by evidence: integrity verification, result-to-claim mapping, and claim auditing that cross-checks manuscript statements against the claim ledger and raw evidence, as well as a five-pass scientific-editing pipeline, mathematical-proof checks, and visual inspection of the rendered PDF. A prototype self-improvement loop records research traces and proposes harness improvements that are adopted only after reviewer approval.","mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/662f43dbe93bb73804fa1606/nFeWKD7V2DRVrzZF_JTDr.png"],"thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2605.03042.png","numComments":10,"upvoted":false,"submittedBy":{"_id":"662f43dbe93bb73804fa1606","avatarUrl":"/avatars/1a3c117283598f566fc59b02e0df2f9e.svg","fullname":"Ruofeng Yang","name":"RuofengYang","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"isAuthorParticipating":true},{"paper":{"id":"2607.19191","authors":[{"_id":"6a60336e7e7f152167e470df","user":{"_id":"6414106ce7d5f817d204e160","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6414106ce7d5f817d204e160/i-AeOvUzm7CIJ5d3ZacjX.png","isPro":false,"fullname":"Frank Jiang","user":"frankjiang","type":"user","name":"frankjiang"},"name":"Fan Jiang","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:56:40.299Z","hidden":false},{"_id":"6a60336e7e7f152167e470e0","name":"Zhaoxu Sun","hidden":false},{"_id":"6a60336e7e7f152167e470e1","name":"Mengchao Wang","hidden":false},{"_id":"6a60336e7e7f152167e470e2","name":"Ziyu Zhu","hidden":false},{"_id":"6a60336e7e7f152167e470e3","name":"Chiyu Wang","hidden":false},{"_id":"6a60336e7e7f152167e470e4","name":"Yunpeng Zhang","hidden":false},{"_id":"6a60336e7e7f152167e470e5","name":"Wenlin Liu","hidden":false},{"_id":"6a60336e7e7f152167e470e6","name":"Yun Wang","hidden":false},{"_id":"6a60336e7e7f152167e470e7","name":"Xue Zheng","hidden":false},{"_id":"6a60336e7e7f152167e470e8","name":"Rui Sun","hidden":false},{"_id":"6a60336e7e7f152167e470e9","name":"Junfeng Ni","hidden":false},{"_id":"6a60336e7e7f152167e470ea","name":"Hongyu Pan","hidden":false},{"_id":"6a60336e7e7f152167e470eb","name":"Zhongxu Sun","hidden":false},{"_id":"6a60336e7e7f152167e470ec","name":"Fei Yu","hidden":false},{"_id":"6a60336e7e7f152167e470ed","name":"Zengye Ge","hidden":false},{"_id":"6a60336e7e7f152167e470ee","name":"Mengmeng Du","hidden":false},{"_id":"6a60336e7e7f152167e470ef","name":"Nianfei Fan","hidden":false},{"_id":"6a60336e7e7f152167e470f0","name":"Mingchao Sun","hidden":false},{"_id":"6a60336e7e7f152167e470f1","name":"Yu Liu","hidden":false},{"_id":"6a60336e7e7f152167e470f2","name":"Yongchang","hidden":false},{"_id":"6a60336e7e7f152167e470f3","name":"Yanqing Zhu","hidden":false},{"_id":"6a60336e7e7f152167e470f4","name":"Jiahang Wang","hidden":false},{"_id":"6a60336e7e7f152167e470f5","name":"Ning Ying","hidden":false},{"_id":"6a60336e7e7f152167e470f6","name":"Yuze Xuan","hidden":false},{"_id":"6a60336e7e7f152167e470f7","name":"Di Yang","hidden":false},{"_id":"6a60336e7e7f152167e470f8","name":"Zhicheng Liu","hidden":false},{"_id":"6a60336e7e7f152167e470f9","name":"Zhe Gao","hidden":false},{"_id":"6a60336e7e7f152167e470fa","name":"Tingbing Xu","hidden":false},{"_id":"6a60336e7e7f152167e470fb","name":"Jiacheng Sui","hidden":false},{"_id":"6a60336e7e7f152167e470fc","name":"Wenjin Yang","hidden":false},{"_id":"6a60336e7e7f152167e470fd","name":"Junnan Lai","hidden":false},{"_id":"6a60336e7e7f152167e470fe","name":"Shufeng Liu","hidden":false},{"_id":"6a60336e7e7f152167e470ff","name":"Yuan Liu","hidden":false},{"_id":"6a60336e7e7f152167e47100","name":"Zheng Zhou","hidden":false},{"_id":"6a60336e7e7f152167e47101","name":"Yingliang Peng","hidden":false},{"_id":"6a60336e7e7f152167e47102","name":"Dawei Cao","hidden":false},{"_id":"6a60336e7e7f152167e47103","name":"Kaifeng Sheng","hidden":false},{"_id":"6a60336e7e7f152167e47104","name":"Yuxiang Cai","hidden":false},{"_id":"6a60336e7e7f152167e47105","name":"Fei Lu","hidden":false},{"_id":"6a60336e7e7f152167e47106","name":"Mu Xu","hidden":false},{"_id":"6a60336e7e7f152167e47107","name":"Ning Guo","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/ew4mZ_fcuneg7qXzDabcH.mp4"],"publishedAt":"2026-07-21T15:26:50.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.","upvotes":311,"discussionId":"6a60336e7e7f152167e47108","projectPage":"https://abot-world.amap.com/","githubRepo":"https://github.com/amap-cvlab/ABot-World","githubRepoAddedBy":"user","ai_summary":"ABot-World-0 is a real-time action-conditioned video world model that uses progressive distillation, long-horizon alignment, and a co-designed streaming stack to enable efficient, long-horizon interactive world generation.","ai_keywords":["action-conditioned video world model","WorldExplorer","VLM-based assessment","ODE distillation","LongForcing","autoregressive drift","reference-character memory","streaming inference stack","VAE decoder","DiT inference","low-bit DiT","WorldRoamBench"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2408,"organization":{"_id":"641415d08900ef6afa2fcb73","name":"acvlab","fullname":"Alibaba AMAP CV Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6414106ce7d5f817d204e160/dfveRtrRy8Xn7QpG684zl.png"}},"publishedAt":"2026-07-21T11:26:50.000Z","title":"ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU","summary":"We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.","mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/ew4mZ_fcuneg7qXzDabcH.mp4"],"thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2607.19191.png","numComments":5,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"organization":{"_id":"641415d08900ef6afa2fcb73","name":"acvlab","fullname":"Alibaba AMAP CV Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6414106ce7d5f817d204e160/dfveRtrRy8Xn7QpG684zl.png"},"isAuthorParticipating":true},{"paper":{"id":"2605.12500","authors":[{"_id":"6a03e33986b054ce2fa40dcc","user":{"_id":"64b4a717aa03b6520839e9b8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b4a717aa03b6520839e9b8/Rt3ERG-6BVEA4hAwOz0_I.jpeg","isPro":false,"fullname":"Haiwen Diao","user":"Paranioar","type":"user","name":"Paranioar"},"name":"Haiwen Diao","status":"claimed_verified","statusLastChangedAt":"2026-05-14T10:58:16.335Z","hidden":false},{"_id":"6a03e33986b054ce2fa40dcd","name":"Penghao Wu","hidden":false},{"_id":"6a03e33986b054ce2fa40dce","name":"Hanming Deng","hidden":false},{"_id":"6a03e33986b054ce2fa40dcf","user":{"_id":"66bc7862aa7cdcb1c31a1efb","avatarUrl":"/avatars/4c2ab907247fe071ff5cdd71c404ca7c.svg","isPro":false,"fullname":"wang jiahao","user":"TokenWang","type":"user","name":"TokenWang"},"name":"Jiahao Wang","status":"claimed_verified","statusLastChangedAt":"2026-05-13T07:44:20.693Z","hidden":false},{"_id":"6a03e33986b054ce2fa40dd0","name":"Shihao Bai","hidden":false},{"_id":"6a03e33986b054ce2fa40dd1","name":"Silei Wu","hidden":false},{"_id":"6a03e33986b054ce2fa40dd2","user":{"_id":"6481764e8af4675862efb22e","avatarUrl":"/avatars/fc2e076bc861693f598a528a068a696e.svg","isPro":false,"fullname":"weichenfan","user":"weepiess2383","type":"user","name":"weepiess2383"},"name":"Weichen Fan","status":"claimed_verified","statusLastChangedAt":"2026-05-28T15:35:16.475Z","hidden":false},{"_id":"6a03e33986b054ce2fa40dd3","name":"Wenjie Ye","hidden":false},{"_id":"6a03e33986b054ce2fa40dd4","user":{"_id":"647d4f1236e109abce409c3b","avatarUrl":"/avatars/d166f5f8be666e96b522a0a0effd21c4.svg","isPro":false,"fullname":"Wenwen Tong","user":"tongww","type":"user","name":"tongww"},"name":"Wenwen Tong","status":"claimed_verified","statusLastChangedAt":"2026-05-14T10:58:12.191Z","hidden":false},{"_id":"6a03e33986b054ce2fa40dd5","name":"Xiangyu Fan","hidden":false},{"_id":"6a03e33986b054ce2fa40dd6","name":"Yan Li","hidden":false},{"_id":"6a03e33986b054ce2fa40dd7","name":"Yubo Wang","hidden":false},{"_id":"6a03e33986b054ce2fa40dd8","name":"Zhijie Cao","hidden":false},{"_id":"6a03e33986b054ce2fa40dd9","user":{"_id":"6583f8dacbb381e788840f06","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6583f8dacbb381e788840f06/tDsK0v8cvYu1vRb2qQRPu.jpeg","isPro":false,"fullname":"zhiqian_joy","user":"zhiqian-joy","type":"user","name":"zhiqian-joy"},"name":"Zhiqian Lin","status":"claimed_verified","statusLastChangedAt":"2026-05-26T07:55:09.717Z","hidden":false},{"_id":"6a03e33986b054ce2fa40dda","name":"Zhitao Yang","hidden":false},{"_id":"6a03e33986b054ce2fa40ddb","user":{"_id":"652d06833b5997ed71ce5c46","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/O_D6bpa5mGxLA7uCjmVCG.jpeg","isPro":false,"fullname":"Zhongang Cai","user":"caizhongang","type":"user","name":"caizhongang"},"name":"Zhongang Cai","status":"claimed_verified","statusLastChangedAt":"2026-05-13T07:44:18.517Z","hidden":false},{"_id":"6a03e33986b054ce2fa40ddc","user":{"_id":"66915a572c1a3a8edcc977b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66915a572c1a3a8edcc977b4/2tANTgj48VQMgCcEcdkwE.jpeg","isPro":false,"fullname":"Yuwei Niu","user":"Yuwei-Niu","type":"user","name":"Yuwei-Niu"},"name":"Yuwei Niu","status":"claimed_verified","statusLastChangedAt":"2026-05-14T10:58:14.215Z","hidden":false},{"_id":"6a03e33986b054ce2fa40ddd","name":"Yue Zhu","hidden":false},{"_id":"6a03e33986b054ce2fa40dde","name":"Bo Liu","hidden":false},{"_id":"6a03e33986b054ce2fa40ddf","name":"Chengguang Lv","hidden":false},{"_id":"6a03e33986b054ce2fa40de0","name":"Haojia Yu","hidden":false},{"_id":"6a03e33986b054ce2fa40de1","user":{"_id":"63f47b5321eb234ab739e91a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f47b5321eb234ab739e91a/vWfFNVtMkHl8gieha5PPd.jpeg","isPro":false,"fullname":"Haozhe Xie","user":"hzxie","type":"user","name":"hzxie"},"name":"Haozhe Xie","status":"claimed_verified","statusLastChangedAt":"2026-05-14T10:58:09.979Z","hidden":false},{"_id":"6a03e33986b054ce2fa40de2","name":"Hongli Wang","hidden":false},{"_id":"6a03e33986b054ce2fa40de3","user":{"_id":"662df5a92b1b529a43a43310","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662df5a92b1b529a43a43310/0zplLZcHue064FuVrH7vl.jpeg","isPro":false,"fullname":"Jianan Fan","user":"muyiyunzi","type":"user","name":"muyiyunzi"},"name":"Jianan Fan","status":"claimed_verified","statusLastChangedAt":"2026-05-13T07:44:14.299Z","hidden":false},{"_id":"6a03e33986b054ce2fa40de4","name":"Jiaqi Li","hidden":false},{"_id":"6a03e33986b054ce2fa40de5","name":"Jiefan Lu","hidden":false},{"_id":"6a03e33986b054ce2fa40de6","name":"Jingcheng Ni","hidden":false},{"_id":"6a03e33986b054ce2fa40de7","user":{"_id":"680b0d0f1173808fedf31f96","avatarUrl":"/avatars/ec920acf1200fc301737529aaabc87db.svg","isPro":false,"fullname":"Junxiang Xu","user":"junxiangjoe","type":"user","name":"junxiangjoe"},"name":"Junxiang Xu","status":"claimed_verified","statusLastChangedAt":"2026-05-18T09:50:30.702Z","hidden":false},{"_id":"6a03e33986b054ce2fa40de8","name":"Kaihuan Liang","hidden":false},{"_id":"6a03e33986b054ce2fa40de9","name":"Lianqiang Shi","hidden":false},{"_id":"6a03e33986b054ce2fa40dea","name":"Linjun Dai","hidden":false},{"_id":"6a03e33986b054ce2fa40deb","name":"Linyan Wang","hidden":false},{"_id":"6a03e33986b054ce2fa40dec","user":{"_id":"647ae5462d27d3541deb70ff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/BK7z2jhmKq3asDuiGmT3L.png","isPro":false,"fullname":"Oscar J Qian","user":"oscarqjh","type":"user","name":"oscarqjh"},"name":"Oscar Qian","status":"claimed_verified","statusLastChangedAt":"2026-05-14T10:58:07.426Z","hidden":false},{"_id":"6a03e33986b054ce2fa40ded","user":{"_id":"63ed9040c5b3c73408642ffb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ed9040c5b3c73408642ffb/-xxR7lULo3TzizwSgl3H5.png","isPro":false,"fullname":"GaoPeng","user":"gaclove","type":"user","name":"gaclove"},"name":"Peng Gao","status":"claimed_verified","statusLastChangedAt":"2026-05-13T07:44:16.488Z","hidden":false},{"_id":"6a03e33986b054ce2fa40dee","name":"Pengfei Liu","hidden":false},{"_id":"6a03e33986b054ce2fa40def","name":"Qingping Sun","hidden":false},{"_id":"6a03e33986b054ce2fa40df0","name":"Rui Shen","hidden":false},{"_id":"6a03e33986b054ce2fa40df1","name":"Ruisi Wang","hidden":false},{"_id":"6a03e33986b054ce2fa40df2","name":"Shengnan Ma","hidden":false},{"_id":"6a03e33986b054ce2fa40df3","name":"Shuang Yang","hidden":false},{"_id":"6a03e33986b054ce2fa40df4","name":"Siyi Xie","hidden":false},{"_id":"6a03e33986b054ce2fa40df5","name":"Siying Li","hidden":false},{"_id":"6a03e33986b054ce2fa40df6","name":"Tianbo Zhong","hidden":false},{"_id":"6a03e33986b054ce2fa40df7","name":"Xiangli Kong","hidden":false},{"_id":"6a03e33986b054ce2fa40df8","user":{"_id":"634f8986d049354d7ee93db2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634f8986d049354d7ee93db2/CKYqG-m2q0Jf2_TpNJwwd.jpeg","isPro":false,"fullname":"Xuanke Shi","user":"shixuanke","type":"user","name":"shixuanke"},"name":"Xuanke Shi","status":"claimed_verified","statusLastChangedAt":"2026-05-22T16:10:04.720Z","hidden":false},{"_id":"6a03e33986b054ce2fa40df9","name":"Yang Gao","hidden":false},{"_id":"6a03e33986b054ce2fa40dfa","name":"Yongqiang Yao","hidden":false},{"_id":"6a03e33986b054ce2fa40dfb","name":"Yves Wang","hidden":false},{"_id":"6a03e33986b054ce2fa40dfc","name":"Zhengqi Bai","hidden":false},{"_id":"6a03e33986b054ce2fa40dfd","name":"Zhengyu Lin","hidden":false},{"_id":"6a03e33986b054ce2fa40dfe","name":"Zixin Yin","hidden":false},{"_id":"6a03e33986b054ce2fa40dff","name":"Wenxiu Sun","hidden":false},{"_id":"6a03e33986b054ce2fa40e00","name":"Ruihao Gong","hidden":false},{"_id":"6a03e33986b054ce2fa40e01","name":"Quan Wang","hidden":false},{"_id":"6a03e33986b054ce2fa40e02","name":"Lewei Lu","hidden":false},{"_id":"6a03e33986b054ce2fa40e03","user":{"_id":"6626a471430a124253f197c8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6626a471430a124253f197c8/uVEk5nnW-bS6-no0rQ7Wh.png","isPro":false,"fullname":"yl-1993","user":"yl-1993","type":"user","name":"yl-1993"},"name":"Lei Yang","status":"claimed_verified","statusLastChangedAt":"2026-05-13T07:44:11.767Z","hidden":false},{"_id":"6a03e33986b054ce2fa40e04","name":"Ziwei Liu","hidden":false},{"_id":"6a03e33986b054ce2fa40e05","name":"Dahua Lin","hidden":false}],"publishedAt":"2026-05-12T00:00:00.000Z","submittedOnDailyAt":"2026-05-13T00:00:00.000Z","title":"SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture","submittedOnDailyBy":{"_id":"64b4a717aa03b6520839e9b8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b4a717aa03b6520839e9b8/Rt3ERG-6BVEA4hAwOz0_I.jpeg","isPro":false,"fullname":"Haiwen Diao","user":"Paranioar","type":"user","name":"Paranioar"},"summary":"Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT, built on dense (8B) and mixture-of-experts (30B-A3B) understanding baselines, respectively. Designed from first principles, they rival top-tier understanding-only VLMs across text understanding, vision-language perception, knowledge reasoning, agentic decision-making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge-intensive any-to-image (X2I) synthesis, complex text-rich infographic generation, and interleaved vision-language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre-/post-training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision-language-action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.","upvotes":195,"discussionId":"6a03e33a86b054ce2fa40e06","githubRepo":"https://github.com/OpenSenseNova/SenseNova-U1","githubRepoAddedBy":"user","ai_summary":"Unified vision-language models treat understanding and generation as integrated processes rather than separate tasks, demonstrating strong performance across multiple multimodal capabilities including image synthesis and action reasoning.","ai_keywords":["vision-language models","multimodal intelligence","unified paradigm","NEO-unify","dense model","mixture-of-experts","vision-language perception","knowledge reasoning","agentic decision-making","spatial intelligence","any-to-image synthesis","text-rich infographic generation","vision-language-action","world model"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":4872,"organization":{"_id":"64f0405f8a4cf3e5e6b38f9c","name":"sensenova","fullname":"SenseNova","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/k66xcOMf4NVbMSFulUjHY.png"}},"publishedAt":"2026-05-11T20:00:00.000Z","title":"SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture","summary":"Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT, built on dense (8B) and mixture-of-experts (30B-A3B) understanding baselines, respectively. Designed from first principles, they rival top-tier understanding-only VLMs across text understanding, vision-language perception, knowledge reasoning, agentic decision-making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge-intensive any-to-image (X2I) synthesis, complex text-rich infographic generation, and interleaved vision-language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre-/post-training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision-language-action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2605.12500.png","numComments":2,"upvoted":false,"submittedBy":{"_id":"64b4a717aa03b6520839e9b8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b4a717aa03b6520839e9b8/Rt3ERG-6BVEA4hAwOz0_I.jpeg","fullname":"Haiwen Diao","name":"Paranioar","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":53,"isUserFollowing":false},"organization":{"_id":"64f0405f8a4cf3e5e6b38f9c","name":"sensenova","fullname":"SenseNova","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/k66xcOMf4NVbMSFulUjHY.png"},"isAuthorParticipating":true},{"paper":{"id":"2608.13552","authors":[{"_id":"6a7e730742823931a1f17557","user":{"_id":"63185a65d369578578fa1ed4","avatarUrl":"/avatars/1cb7d593d2f430ebf22b0e80e4dea9d0.svg","isPro":false,"fullname":"Kaixin Ding","user":"jocelynd","type":"user","name":"jocelynd"},"name":"Kaixin Ding","status":"claimed_verified","statusLastChangedAt":"2026-08-14T08:45:04.611Z","hidden":false},{"_id":"6a7e730742823931a1f17558","name":"Xi Chen","hidden":false},{"_id":"6a7e730742823931a1f17559","name":"Minghong Cai","hidden":false},{"_id":"6a7e730742823931a1f1755a","name":"Zhiyuan Xu","hidden":false},{"_id":"6a7e730742823931a1f1755b","name":"Yiyang Wang","hidden":false},{"_id":"6a7e730742823931a1f1755c","name":"Yuxiang Lu","hidden":false},{"_id":"6a7e730742823931a1f1755d","name":"Junyi Li","hidden":false},{"_id":"6a7e730742823931a1f1755e","name":"Shuyang Chen","hidden":false},{"_id":"6a7e730742823931a1f1755f","name":"Yuan Gao","hidden":false},{"_id":"6a7e730742823931a1f17560","name":"Xin Tao","hidden":false},{"_id":"6a7e730742823931a1f17561","name":"Pengfei Wan","hidden":false},{"_id":"6a7e730742823931a1f17562","name":"Hengshuang Zhao","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives","submittedOnDailyBy":{"_id":"64f94370c3c12b377cc51086","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f94370c3c12b377cc51086/6CXcHhqAoykqXcShqM8Rd.jpeg","isPro":false,"fullname":"Minghong Cai","user":"onevfall","type":"user","name":"onevfall"},"summary":"Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.","upvotes":43,"discussionId":"6a7e730842823931a1f17563","projectPage":"https://kxding.github.io/project/PlayWorld/","githubRepo":"https://github.com/kxding/PlayWorld","githubRepoAddedBy":"user","ai_summary":"PlayWorld benchmarks interactive video world models by using multi-modal agents to pursue long-horizon objectives, evaluating geometry consistency, interaction fidelity, and state evolution.","ai_keywords":["world models","multi-modal Agent Players","PlayWorld","geometry consistency","interaction fidelity","out-of-sight evolution","insight evolution"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":54,"organization":{"_id":"67ea9ecfc234715db8dbf339","name":"hkuhk","fullname":"The University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ea9e8d2d95c10a0da11b0c/FNnR4M7YqKRuG43N5771B.png"}},"publishedAt":"2026-08-12T20:00:00.000Z","title":"PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives","summary":"Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2608.13552.png","numComments":2,"upvoted":false,"submittedBy":{"_id":"64f94370c3c12b377cc51086","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f94370c3c12b377cc51086/6CXcHhqAoykqXcShqM8Rd.jpeg","fullname":"Minghong Cai","name":"onevfall","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"organization":{"_id":"67ea9ecfc234715db8dbf339","name":"hkuhk","fullname":"The University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ea9e8d2d95c10a0da11b0c/FNnR4M7YqKRuG43N5771B.png"},"isAuthorParticipating":false},{"paper":{"id":"2501.13956","authors":[{"_id":"67be29ff8c5a19f7c08ab4b2","name":"Preston Rasmussen","hidden":false},{"_id":"67be29ff8c5a19f7c08ab4b3","user":{"_id":"66135780174b378a723f2ae9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/bIzz1wzTD2wn7sokzUc4u.png","isPro":false,"fullname":"Paul Paliychuk","user":"paulzep","type":"user","name":"paulzep"},"name":"Pavlo Paliychuk","status":"extracted_confirmed","statusLastChangedAt":"2025-02-25T21:04:47.368Z","hidden":false},{"_id":"67be29ff8c5a19f7c08ab4b4","name":"Travis Beauvais","hidden":false},{"_id":"67be29ff8c5a19f7c08ab4b5","name":"Jack Ryan","hidden":false},{"_id":"67be29ff8c5a19f7c08ab4b6","user":{"_id":"630a77ccb70f9eb690fb9285","avatarUrl":"/avatars/62a23c838dacd39cca55e95c5788e601.svg","isPro":false,"fullname":"Daniel Chalef","user":"danielchalef","type":"user","name":"danielchalef"},"name":"Daniel Chalef","status":"extracted_confirmed","statusLastChangedAt":"2025-02-25T20:59:34.275Z","hidden":false}],"publishedAt":"2025-01-20T16:52:48.000Z","title":"Zep: A Temporal Knowledge Graph Architecture for Agent Memory","summary":"We introduce Zep, a novel memory layer service for AI agents that outperforms\nthe current state-of-the-art system, MemGPT, in the Deep Memory Retrieval (DMR)\nbenchmark. Additionally, Zep excels in more comprehensive and challenging\nevaluations than DMR that better reflect real-world enterprise use cases. While\nexisting retrieval-augmented generation (RAG) frameworks for large language\nmodel (LLM)-based agents are limited to static document retrieval, enterprise\napplications demand dynamic knowledge integration from diverse sources\nincluding ongoing conversations and business data. Zep addresses this\nfundamental limitation through its core component Graphiti -- a\ntemporally-aware knowledge graph engine that dynamically synthesizes both\nunstructured conversational data and structured business data while maintaining\nhistorical relationships. In the DMR benchmark, which the MemGPT team\nestablished as their primary evaluation metric, Zep demonstrates superior\nperformance (94.8% vs 93.4%). Beyond DMR, Zep's capabilities are further\nvalidated through the more challenging LongMemEval benchmark, which better\nreflects enterprise use cases through complex temporal reasoning tasks. In this\nevaluation, Zep achieves substantial results with accuracy improvements of up\nto 18.5% while simultaneously reducing response latency by 90% compared to\nbaseline implementations. These results are particularly pronounced in\nenterprise-critical tasks such as cross-session information synthesis and\nlong-term context maintenance, demonstrating Zep's effectiveness for deployment\nin real-world applications.","upvotes":16,"discussionId":"67be29ff8c5a19f7c08ab4ed","githubRepo":"https://github.com/getzep/graphiti","githubRepoAddedBy":"auto","ai_summary":"Zep, a memory layer service, outperforms MemGPT in the DMR benchmark and LongMemEval by excelling in dynamic knowledge integration and temporal reasoning, critical for enterprise use cases.","ai_keywords":["memory layer service","Deep Memory Retrieval (DMR)","MemGPT","retrieval-augmented generation (RAG)","large language model (LLM)","Graphiti","knowledge graph engine","temporally-aware","unstructured conversational data","structured business data","LongMemEval","cross-session information synthesis","long-term context maintenance"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":29985},"publishedAt":"2025-01-20T11:52:48.000Z","title":"Zep: A Temporal Knowledge Graph Architecture for Agent Memory","summary":"We introduce Zep, a novel memory layer service for AI agents that outperforms\nthe current state-of-the-art system, MemGPT, in the Deep Memory Retrieval (DMR)\nbenchmark. Additionally, Zep excels in more comprehensive and challenging\nevaluations than DMR that better reflect real-world enterprise use cases. While\nexisting retrieval-augmented generation (RAG) frameworks for large language\nmodel (LLM)-based agents are limited to static document retrieval, enterprise\napplications demand dynamic knowledge integration from diverse sources\nincluding ongoing conversations and business data. Zep addresses this\nfundamental limitation through its core component Graphiti -- a\ntemporally-aware knowledge graph engine that dynamically synthesizes both\nunstructured conversational data and structured business data while maintaining\nhistorical relationships. In the DMR benchmark, which the MemGPT team\nestablished as their primary evaluation metric, Zep demonstrates superior\nperformance (94.8% vs 93.4%). Beyond DMR, Zep's capabilities are further\nvalidated through the more challenging LongMemEval benchmark, which better\nreflects enterprise use cases through complex temporal reasoning tasks. In this\nevaluation, Zep achieves substantial results with accuracy improvements of up\nto 18.5% while simultaneously reducing response latency by 90% compared to\nbaseline implementations. These results are particularly pronounced in\nenterprise-critical tasks such as cross-session information synthesis and\nlong-term context maintenance, demonstrating Zep's effectiveness for deployment\nin real-world applications.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2501.13956.png","numComments":0,"upvoted":false,"isAuthorParticipating":false},{"paper":{"id":"2606.03748","authors":[{"_id":"6a1f9080e292c1c78ecb132f","name":"Glenn Jocher","hidden":false},{"_id":"6a1f9080e292c1c78ecb1330","name":"Jing Qiu","hidden":false},{"_id":"6a1f9080e292c1c78ecb1331","name":"Mengyu Liu","hidden":false},{"_id":"6a1f9080e292c1c78ecb1332","name":"Shuai Lyu","hidden":false},{"_id":"6a1f9080e292c1c78ecb1333","user":{"_id":"617e7dbd129c9e67703ffe62","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/617e7dbd129c9e67703ffe62/nJD_rI351t2c2QN_BpUBz.jpeg","isPro":false,"fullname":"Fatih C. Akyon","user":"fcakyon","type":"user","name":"fcakyon"},"name":"Fatih Cagatay Akyon","status":"claimed_verified","statusLastChangedAt":"2026-06-09T12:47:21.531Z","hidden":false},{"_id":"6a1f9080e292c1c78ecb1334","name":"Muhammet Esat Kalfaoglu","hidden":false}],"publishedAt":"2026-06-02T00:00:00.000Z","submittedOnDailyAt":"2026-06-03T00:00:00.000Z","title":"Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models","submittedOnDailyBy":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","isPro":false,"fullname":"Niels Rogge","user":"nielsr","type":"user","name":"nielsr"},"summary":"Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware. The YOLO family has become widely deployed for this reason, yet most YOLO detectors still rely on non-maximum suppression at inference, carry heavy detection heads due to Distribution Focal Loss, require long training schedules, and can leave the smallest objects without positive label assignments. We present Ultralytics YOLO26, a unified real-time vision model family that addresses these limitations through coordinated architecture and training advances. YOLO26 uses a dual-head design for native NMS-free end-to-end inference and removes DFL entirely, yielding a lighter head with unconstrained regression range. Its training pipeline combines MuSGD, a hybrid Muon-SGD optimizer adapted from large language model training; Progressive Loss, which shifts supervision toward the inference-time head; and STAL, a label assignment strategy that guarantees positive coverage for small objects. Beyond detection, YOLO26 introduces task-specific head and loss designs for instance segmentation, pose estimation, and oriented detection, producing consistent gains across tasks and scales. The family spans five scales (n/s/m/l/x) and supports detection, instance segmentation, pose estimation, classification, and oriented detection in a single pipeline, with an open-vocabulary extension, YOLOE-26, for text-, visual-, and prompt-free inference. Across all scales, YOLO26 achieves 40.9-57.5 mAP on COCO at 1.7-11.8 ms T4 TensorRT latency, advancing the accuracy-latency Pareto front over prior real-time detectors, while YOLOE-26x reaches 40.6 AP on LVIS minival under text prompting. Code and models are available at https://github.com/ultralytics/ultralytics.","upvotes":22,"discussionId":"6a1f9080e292c1c78ecb1335","projectPage":"https://docs.ultralytics.com/models/yolo26","githubRepo":"https://github.com/ultralytics/ultralytics","githubRepoAddedBy":"user","ai_summary":"YOLO26 addresses real-time vision challenges through a unified model family with NMS-free inference, improved training strategies, and multi-task capabilities spanning detection, segmentation, and pose estimation.","ai_keywords":["YOLO","non-maximum suppression","Distribution Focal Loss","MuSGD","hybrid Muon-SGD optimizer","Progressive Loss","STAL","instance segmentation","pose estimation","oriented detection","open-vocabulary extension","TensorRT latency","mAP","COCO","LVIS"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":60682,"organization":{"_id":"65ba54e4b6b3191a83624cee","name":"Ultralytics","fullname":"Ultralytics","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60ad037a306d6873ec42d537/iusy6Ia8JeCK-ui4z-kjE.png"}},"publishedAt":"2026-06-01T20:00:00.000Z","title":"Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models","summary":"Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware. The YOLO family has become widely deployed for this reason, yet most YOLO detectors still rely on non-maximum suppression at inference, carry heavy detection heads due to Distribution Focal Loss, require long training schedules, and can leave the smallest objects without positive label assignments. We present Ultralytics YOLO26, a unified real-time vision model family that addresses these limitations through coordinated architecture and training advances. YOLO26 uses a dual-head design for native NMS-free end-to-end inference and removes DFL entirely, yielding a lighter head with unconstrained regression range. Its training pipeline combines MuSGD, a hybrid Muon-SGD optimizer adapted from large language model training; Progressive Loss, which shifts supervision toward the inference-time head; and STAL, a label assignment strategy that guarantees positive coverage for small objects. Beyond detection, YOLO26 introduces task-specific head and loss designs for instance segmentation, pose estimation, and oriented detection, producing consistent gains across tasks and scales. The family spans five scales (n/s/m/l/x) and supports detection, instance segmentation, pose estimation, classification, and oriented detection in a single pipeline, with an open-vocabulary extension, YOLOE-26, for text-, visual-, and prompt-free inference. Across all scales, YOLO26 achieves 40.9-57.5 mAP on COCO at 1.7-11.8 ms T4 TensorRT latency, advancing the accuracy-latency Pareto front over prior real-time detectors, while YOLOE-26x reaches 40.6 AP on LVIS minival under text prompting. Code and models are available at https://github.com/ultralytics/ultralytics.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2606.03748.png","numComments":1,"upvoted":false,"submittedBy":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","fullname":"Niels Rogge","name":"nielsr","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":1278,"isUserFollowing":false},"organization":{"_id":"65ba54e4b6b3191a83624cee","name":"Ultralytics","fullname":"Ultralytics","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60ad037a306d6873ec42d537/iusy6Ia8JeCK-ui4z-kjE.png"},"isAuthorParticipating":false},{"paper":{"id":"2512.16676","authors":[{"_id":"6949026334f46eaf46cbb3d1","user":{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","isPro":true,"fullname":"Hao Liang","user":"lhpku20010120","type":"user","name":"lhpku20010120"},"name":"Hao Liang","status":"claimed_verified","statusLastChangedAt":"2026-04-03T09:09:52.246Z","hidden":false},{"_id":"6949026334f46eaf46cbb3d2","user":{"_id":"65099d08f37afbab0d3fb268","avatarUrl":"/avatars/cef45b7c6b7c90bbef341a39a9bb51be.svg","isPro":false,"fullname":"Xiaochen Ma","user":"Sunnyhaze","type":"user","name":"Sunnyhaze"},"name":"Xiaochen Ma","status":"claimed_verified","statusLastChangedAt":"2025-12-25T20:51:26.454Z","hidden":false},{"_id":"6949026334f46eaf46cbb3d3","name":"Zhou Liu","hidden":false},{"_id":"6949026334f46eaf46cbb3d4","user":{"_id":"65d3149f0545ab7c11568b2b","avatarUrl":"/avatars/02f36d8320b8331849e7d7f40e509ca8.svg","isPro":false,"fullname":"Zhen Hao Wong","user":"aaron1141","type":"user","name":"aaron1141"},"name":"Zhen Hao Wong","status":"claimed_verified","statusLastChangedAt":"2025-12-25T20:51:24.004Z","hidden":false},{"_id":"6949026334f46eaf46cbb3d5","name":"Zhengyang Zhao","hidden":false},{"_id":"6949026334f46eaf46cbb3d6","name":"Zimo Meng","hidden":false},{"_id":"6949026334f46eaf46cbb3d7","user":{"_id":"65b7098af327f1f4e315294d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b7098af327f1f4e315294d/38JtHebJ_XYWciMA15vJ3.jpeg","isPro":false,"fullname":"Runming He","user":"blackBOX25I47","type":"user","name":"blackBOX25I47"},"name":"Runming He","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.075Z","hidden":false},{"_id":"6949026334f46eaf46cbb3d8","name":"Chengyu Shen","hidden":false},{"_id":"6949026334f46eaf46cbb3d9","name":"Qifeng Cai","hidden":false},{"_id":"6949026334f46eaf46cbb3da","name":"Zhaoyang Han","hidden":false},{"_id":"6949026334f46eaf46cbb3db","name":"Meiyi Qiang","hidden":false},{"_id":"6949026334f46eaf46cbb3dc","name":"Yalin Feng","hidden":false},{"_id":"6949026334f46eaf46cbb3dd","name":"Tianyi Bai","hidden":false},{"_id":"6949026334f46eaf46cbb3de","user":{"_id":"658a907a15b65eb9ba2f52e8","avatarUrl":"/avatars/64575a37a6e449ad525877a70eafa1be.svg","isPro":false,"fullname":"zewei pan","user":"pzp5700","type":"user","name":"pzp5700"},"name":"Zewei Pan","status":"claimed_verified","statusLastChangedAt":"2025-12-29T14:25:06.525Z","hidden":false},{"_id":"6949026334f46eaf46cbb3df","name":"Ziyi Guo","hidden":false},{"_id":"6949026334f46eaf46cbb3e0","name":"Yizhen Jiang","hidden":false},{"_id":"6949026334f46eaf46cbb3e1","name":"Jingwen Deng","hidden":false},{"_id":"6949026334f46eaf46cbb3e2","name":"Qijie You","hidden":false},{"_id":"6949026334f46eaf46cbb3e3","name":"Peichao Lai","hidden":false},{"_id":"6949026334f46eaf46cbb3e4","name":"Tianyu Guo","hidden":false},{"_id":"6949026334f46eaf46cbb3e5","name":"Chi Hsu Tsai","hidden":false},{"_id":"6949026334f46eaf46cbb3e6","name":"Hengyi Feng","hidden":false},{"_id":"6949026334f46eaf46cbb3e7","name":"Rui Hu","hidden":false},{"_id":"6949026334f46eaf46cbb3e8","name":"Wenkai Yu","hidden":false},{"_id":"6949026334f46eaf46cbb3e9","name":"Junbo Niu","hidden":false},{"_id":"6949026334f46eaf46cbb3ea","user":{"_id":"6671214c92412fd4640714eb","avatarUrl":"/avatars/48fa84e7bc3bb92ad0192aa26b32de10.svg","isPro":false,"fullname":"Bohan Zeng","user":"zbhpku","type":"user","name":"zbhpku"},"name":"Bohan Zeng","status":"claimed_verified","statusLastChangedAt":"2026-04-07T08:49:41.764Z","hidden":false},{"_id":"6949026334f46eaf46cbb3eb","name":"Ruichuan An","hidden":false},{"_id":"6949026334f46eaf46cbb3ec","name":"Lu Ma","hidden":false},{"_id":"6949026334f46eaf46cbb3ed","name":"Jihao Huang","hidden":false},{"_id":"6949026334f46eaf46cbb3ee","user":{"_id":"642fef28a043f0ac7defa8a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/RwOEkuj3fOnOA54tGR7Ea.png","isPro":false,"fullname":"Yaowei Zheng","user":"hiyouga","type":"user","name":"hiyouga"},"name":"Yaowei Zheng","status":"claimed_verified","statusLastChangedAt":"2025-12-31T20:57:58.594Z","hidden":false},{"_id":"6949026334f46eaf46cbb3ef","name":"Conghui He","hidden":false},{"_id":"6949026334f46eaf46cbb3f0","name":"Linpeng Tang","hidden":false},{"_id":"6949026334f46eaf46cbb3f1","name":"Bin Cui","hidden":false},{"_id":"6949026334f46eaf46cbb3f2","name":"Weinan E","hidden":false},{"_id":"6949026334f46eaf46cbb3f3","name":"Wentao Zhang","hidden":false}],"publishedAt":"2025-12-18T15:46:15.000Z","submittedOnDailyAt":"2025-12-23T00:00:00.000Z","title":"DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI","submittedOnDailyBy":{"_id":"6671214c92412fd4640714eb","avatarUrl":"/avatars/48fa84e7bc3bb92ad0192aa26b32de10.svg","isPro":false,"fullname":"Bohan Zeng","user":"zbhpku","type":"user","name":"zbhpku"},"summary":"The rapidly growing demand for high-quality data in Large Language Models (LLMs) has intensified the need for scalable, reliable, and semantically rich data preparation pipelines. However, current practices remain dominated by ad-hoc scripts and loosely specified workflows, which lack principled abstractions, hinder reproducibility, and offer limited support for model-in-the-loop data generation. To address these challenges, we present DataFlow, a unified and extensible LLM-driven data preparation framework. DataFlow is designed with system-level abstractions that enable modular, reusable, and composable data transformations, and provides a PyTorch-style pipeline construction API for building debuggable and optimizable dataflows. The framework consists of nearly 200 reusable operators and six domain-general pipelines spanning text, mathematical reasoning, code, Text-to-SQL, agentic RAG, and large-scale knowledge extraction. To further improve usability, we introduce DataFlow-Agent, which automatically translates natural-language specifications into executable pipelines via operator synthesis, pipeline planning, and iterative verification. Across six representative use cases, DataFlow consistently improves downstream LLM performance. Our math, code, and text pipelines outperform curated human datasets and specialized synthetic baselines, achieving up to +3\\% execution accuracy in Text-to-SQL over SynSQL, +7\\% average improvements on code benchmarks, and 1--3 point gains on MATH, GSM8K, and AIME. Moreover, a unified 10K-sample dataset produced by DataFlow enables base models to surpass counterparts trained on 1M Infinity-Instruct data. These results demonstrate that DataFlow provides a practical and high-performance substrate for reliable, reproducible, and scalable LLM data preparation, and establishes a system-level foundation for future data-centric AI development.","upvotes":226,"discussionId":"6949026334f46eaf46cbb3f4","projectPage":"https://github.com/OpenDCAI/DataFlow","githubRepo":"https://github.com/OpenDCAI/DataFlow","githubRepoAddedBy":"user","ai_summary":"DataFlow is an LLM-driven data preparation framework that enhances data quality and reproducibility for various tasks, improving LLM performance with automatically generated pipelines.","ai_keywords":["DataFlow","Large Language Models (LLMs)","data preparation pipelines","system-level abstractions","PyTorch-style pipeline construction API","reusable operators","domain-general pipelines","Text-to-SQL","agentic RAG","large-scale knowledge extraction","DataFlow-Agent","operator synthesis","pipeline planning","iterative verification"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":7500,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"}},"publishedAt":"2025-12-18T10:46:15.000Z","title":"DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI","summary":"The rapidly growing demand for high-quality data in Large Language Models (LLMs) has intensified the need for scalable, reliable, and semantically rich data preparation pipelines. However, current practices remain dominated by ad-hoc scripts and loosely specified workflows, which lack principled abstractions, hinder reproducibility, and offer limited support for model-in-the-loop data generation. To address these challenges, we present DataFlow, a unified and extensible LLM-driven data preparation framework. DataFlow is designed with system-level abstractions that enable modular, reusable, and composable data transformations, and provides a PyTorch-style pipeline construction API for building debuggable and optimizable dataflows. The framework consists of nearly 200 reusable operators and six domain-general pipelines spanning text, mathematical reasoning, code, Text-to-SQL, agentic RAG, and large-scale knowledge extraction. To further improve usability, we introduce DataFlow-Agent, which automatically translates natural-language specifications into executable pipelines via operator synthesis, pipeline planning, and iterative verification. Across six representative use cases, DataFlow consistently improves downstream LLM performance. Our math, code, and text pipelines outperform curated human datasets and specialized synthetic baselines, achieving up to +3\\% execution accuracy in Text-to-SQL over SynSQL, +7\\% average improvements on code benchmarks, and 1--3 point gains on MATH, GSM8K, and AIME. Moreover, a unified 10K-sample dataset produced by DataFlow enables base models to surpass counterparts trained on 1M Infinity-Instruct data. These results demonstrate that DataFlow provides a practical and high-performance substrate for reliable, reproducible, and scalable LLM data preparation, and establishes a system-level foundation for future data-centric AI development.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2512.16676.png","numComments":4,"upvoted":false,"submittedBy":{"_id":"6671214c92412fd4640714eb","avatarUrl":"/avatars/48fa84e7bc3bb92ad0192aa26b32de10.svg","fullname":"Bohan Zeng","name":"zbhpku","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":13,"isUserFollowing":false},"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"},"isAuthorParticipating":true},{"paper":{"id":"2607.11562","authors":[{"_id":"6a55ddf2a9d74d6e65bbd889","name":"Yuliang Liu","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd88a","name":"Zhang Li","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd88b","user":{"_id":"641a955891e3376a057b54b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/_Pecp1aYjbjr71x9yBNA6.jpeg","isPro":false,"fullname":"Ziyang Zhang","user":"zenosai","type":"user","name":"zenosai"},"name":"Ziyang Zhang","status":"claimed_verified","statusLastChangedAt":"2026-07-20T16:04:58.022Z","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd88c","name":"Shuo Zhang","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd88d","name":"Qiang Liu","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd88e","name":"Jiajun Song","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd88f","name":"Zidun Guo","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd890","name":"Xinhan Wang","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd891","name":"Handong Zheng","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd892","name":"Yang Liu","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd893","name":"Dongliang Luo","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd894","name":"Zhiyin Ma","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd895","name":"Jiarui Zhang","hidden":false},{"_id":"6a55ddf2a9d74d6e65bbd896","name":"Xiang Bai","hidden":false}],"publishedAt":"2026-07-13T00:00:00.000Z","submittedOnDailyAt":"2026-07-15T00:00:00.000Z","title":"MonkeyOCRv2: A Visual-Text Foundation Model for Document AI","submittedOnDailyBy":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","isPro":false,"fullname":"Niels Rogge","user":"nielsr","type":"user","name":"nielsr"},"summary":"Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11times smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.","upvotes":78,"discussionId":"6a55ddf2a9d74d6e65bbd89d","githubRepo":"https://github.com/Yuliang-Liu/MonkeyOCRv2","githubRepoAddedBy":"user","ai_summary":"MonkeyOCRv2 is a document-oriented visual-text pretrained encoder that improves performance across document analysis and multimodal document understanding tasks while using a much smaller vision backbone.","ai_keywords":["visual-text pretraining","document-image pretraining","image-to-text generation","pixel-level document reconstruction","document analysis","multimodal large language models","document parsing","document understanding"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":923,"organization":{"_id":"67bb321c6792208f4a9d5fa7","name":"VLRLab-OCR","fullname":"VLRLab-OCR","avatar":"https://www.gravatar.com/avatar/42942639a60077803d0ad20d6351f06b?d=retro&size=100"}},"publishedAt":"2026-07-12T20:00:00.000Z","title":"MonkeyOCRv2: A Visual-Text Foundation Model for Document AI","summary":"Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11times smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2607.11562.png","numComments":2,"upvoted":false,"submittedBy":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","fullname":"Niels Rogge","name":"nielsr","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":1278,"isUserFollowing":false},"organization":{"_id":"67bb321c6792208f4a9d5fa7","name":"VLRLab-OCR","fullname":"VLRLab-OCR","avatar":"https://www.gravatar.com/avatar/42942639a60077803d0ad20d6351f06b?d=retro&size=100"},"isAuthorParticipating":false},{"paper":{"id":"2605.23904","authors":[{"_id":"6a13aad74d9e8d8602d201a8","user":{"_id":"62d18eb81e36881a57f29bf4","avatarUrl":"/avatars/104851421b4ee9641daaf15942fa7ea1.svg","isPro":false,"fullname":"Yif Yang","user":"Yif29","type":"user","name":"Yif29"},"name":"Yifan Yang","status":"claimed_verified","statusLastChangedAt":"2026-05-25T15:12:31.312Z","hidden":false},{"_id":"6a13aad74d9e8d8602d201a9","user":{"_id":"660691330be1fbe3b9e4c33d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660691330be1fbe3b9e4c33d/TxrDFH_cRu3AlpMC3xmhv.jpeg","isPro":false,"fullname":"ZiYang Gong","user":"Cusyoung","type":"user","name":"Cusyoung"},"name":"Ziyang Gong","status":"claimed_verified","statusLastChangedAt":"2026-05-25T15:12:29.154Z","hidden":false},{"_id":"6a13aad74d9e8d8602d201aa","name":"Weiquan Huang","hidden":false},{"_id":"6a13aad74d9e8d8602d201ab","name":"Qihao Yang","hidden":false},{"_id":"6a13aad74d9e8d8602d201ac","name":"Ziwei Zhou","hidden":false},{"_id":"6a13aad74d9e8d8602d201ad","user":{"_id":"662b6a8f0b7f23f3c000559e","avatarUrl":"/avatars/0a5b4e09ac9a8e40342131319ff32b29.svg","isPro":false,"fullname":"Zisu Huang","user":"zisuh","type":"user","name":"zisuh"},"name":"Zisu Huang","status":"claimed_verified","statusLastChangedAt":"2026-05-25T15:12:25.187Z","hidden":false},{"_id":"6a13aad74d9e8d8602d201ae","name":"Yan Li","hidden":false},{"_id":"6a13aad74d9e8d8602d201af","name":"Xuemei Gao","hidden":false},{"_id":"6a13aad74d9e8d8602d201b0","user":{"_id":"65115c00a588fdb36558b673","avatarUrl":"/avatars/1f36263dc4bfaf696a4aa959a6aab1e1.svg","isPro":false,"fullname":"Qi Dai","user":"daiqi","type":"user","name":"daiqi"},"name":"Qi Dai","status":"claimed_verified","statusLastChangedAt":"2026-07-20T15:44:38.713Z","hidden":false},{"_id":"6a13aad74d9e8d8602d201b1","name":"Bei Liu","hidden":false},{"_id":"6a13aad74d9e8d8602d201b2","name":"Kai Qiu","hidden":false},{"_id":"6a13aad74d9e8d8602d201b3","name":"Yuqing Yang","hidden":false},{"_id":"6a13aad74d9e8d8602d201b4","name":"Dongdong Chen","hidden":false},{"_id":"6a13aad74d9e8d8602d201b5","name":"Xue Yang","hidden":false},{"_id":"6a13aad74d9e8d8602d201b6","name":"Chong Luo","hidden":false}],"publishedAt":"2026-05-22T00:00:00.000Z","submittedOnDailyAt":"2026-05-25T00:00:00.000Z","title":"SkillOpt: Executive Strategy for Self-Evolving Agent Skills","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point under feedback. We argue the skill should instead be trained as the external state of a frozen agent, with the same discipline that makes weight-space optimization reproducible. SkillOpt is, to our knowledge, the first systematic controllable text-space optimizer for agent skills: a separate optimizer model turns scored rollouts into bounded add/delete/replace edits on a single skill document, and an edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, rejected-edit buffer, and epoch-wise slow/meta update make skill training stable while adding zero inference-time model calls at deployment. Across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex, Claude Code), SkillOpt is best or tied on all 52 evaluated (model, benchmark, harness) cells and beats every per-cell competitor among human, one-shot LLM, Trace2Skill, TextGrad, GEPA, and EvoSkill skills. On GPT-5.5 it lifts the average no-skill accuracy by +23.5 points in direct chat, by +24.8 inside the Codex agentic loop, and by +19.1 inside Claude Code. Transfer experiments further show that optimized skill artifacts retain value when moved across model scales, between Codex and Claude Code execution environments, and to a nearby math benchmark without further optimization.","upvotes":264,"discussionId":"6a13aad74d9e8d8602d201b7","projectPage":"https://microsoft.github.io/SkillOpt/","githubRepo":"https://github.com/microsoft/SkillOpt","githubRepoAddedBy":"user","ai_summary":"SkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.","ai_keywords":["agent skills","skill training","text-space optimizer","rollouts","add/delete/replace edits","validation score","textual learning-rate budget","rejected-edit buffer","epoch-wise slow/meta update","skill optimization","transfer experiments","agent state","reproducible optimization"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":16067,"organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"}},"publishedAt":"2026-05-21T20:00:00.000Z","title":"SkillOpt: Executive Strategy for Self-Evolving Agent Skills","summary":"Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point under feedback. We argue the skill should instead be trained as the external state of a frozen agent, with the same discipline that makes weight-space optimization reproducible. SkillOpt is, to our knowledge, the first systematic controllable text-space optimizer for agent skills: a separate optimizer model turns scored rollouts into bounded add/delete/replace edits on a single skill document, and an edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, rejected-edit buffer, and epoch-wise slow/meta update make skill training stable while adding zero inference-time model calls at deployment. Across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex, Claude Code), SkillOpt is best or tied on all 52 evaluated (model, benchmark, harness) cells and beats every per-cell competitor among human, one-shot LLM, Trace2Skill, TextGrad, GEPA, and EvoSkill skills. On GPT-5.5 it lifts the average no-skill accuracy by +23.5 points in direct chat, by +24.8 inside the Codex agentic loop, and by +19.1 inside Claude Code. Transfer experiments further show that optimized skill artifacts retain value when moved across model scales, between Codex and Claude Code execution environments, and to a nearby math benchmark without further optimization.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2605.23904.png","numComments":5,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"},"isAuthorParticipating":true},{"paper":{"id":"2608.11924","authors":[{"_id":"6a7d23410ac8bee77474ee17","name":"Zhuoyang Qian","hidden":false},{"_id":"6a7d23410ac8bee77474ee18","name":"Biao Wu","hidden":false},{"_id":"6a7d23410ac8bee77474ee19","name":"Yiran Wang","hidden":false},{"_id":"6a7d23410ac8bee77474ee1a","name":"Chris D Yan","hidden":false},{"_id":"6a7d23410ac8bee77474ee1b","name":"Desan Dai","hidden":false},{"_id":"6a7d23410ac8bee77474ee1c","name":"Liangwei Zheng","hidden":false},{"_id":"6a7d23410ac8bee77474ee1d","name":"Jin Jiang","hidden":false},{"_id":"6a7d23410ac8bee77474ee1e","name":"Junsheng Zhang","hidden":false},{"_id":"6a7d23410ac8bee77474ee1f","name":"Wenhao Wang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/62b32a4429a410b7f6b06710/pXhDS4-1KDERQMZFpgMo3.png"],"publishedAt":"2026-08-12T00:00:00.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill","submittedOnDailyBy":{"_id":"62b32a4429a410b7f6b06710","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62b32a4429a410b7f6b06710/VzgvmnlYZWuifZTkIkCxy.jpeg","isPro":false,"fullname":"Wenhao Wang","user":"WenhaoWang","type":"user","name":"WenhaoWang"},"summary":"Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service. Spark-to-Paper separates model-based judgment from deterministic operations that can be directly executed and checked. It further separates experiment planning from reporting, so that required evidence is specified before results are observed and manuscript claims are revised according to measured outcomes. To improve reliability over long research trajectories, the system combines deterministic integrity checks with self-critique and bounds a failure mode we call the Self-Refutation Loop, in which repeated experiments continue to reject the original research objective. Spark-to-Paper also produces editable vector figures through programmatic plotting for experimental results and code-based reconstruction for generated method diagrams. Across eight controlled research topics, Spark-to-Paper achieves 99.5% citation validity and 96.4% figure editability. A controlled ablation increases fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, while adversarial review achieves 74% precision. The full system uses 11.9M tokens, costs $8.1 per manuscript, and requires 3.2 hours on average. These results show that end-to-end research paper generation can be implemented as a lightweight, composable workflow inside existing coding assistants while keeping experimental evidence central to how claims are accepted, revised, or abandoned.","upvotes":278,"discussionId":"6a7d23420ac8bee77474ee20","projectPage":"https://spark-to-paper-skills.github.io/spark-to-paper-skills/","githubRepo":"https://github.com/Spark-To-Paper-Skills/spark-to-paper-skills","githubRepoAddedBy":"user","ai_summary":"Spark-to-Paper is a lightweight, composable workflow inside coding assistants that generates research papers by separating planning from reporting, enforcing evidence-based claim revision, and using integrity checks to reduce fabrication.","ai_keywords":["self-refutation loop","citation validity","figure editability","fabrication detection","adversarial review","composable skills","integrity checks","self-critique"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":619},"publishedAt":"2026-08-11T20:00:00.000Z","title":"Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill","summary":"Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service. Spark-to-Paper separates model-based judgment from deterministic operations that can be directly executed and checked. It further separates experiment planning from reporting, so that required evidence is specified before results are observed and manuscript claims are revised according to measured outcomes. To improve reliability over long research trajectories, the system combines deterministic integrity checks with self-critique and bounds a failure mode we call the Self-Refutation Loop, in which repeated experiments continue to reject the original research objective. Spark-to-Paper also produces editable vector figures through programmatic plotting for experimental results and code-based reconstruction for generated method diagrams. Across eight controlled research topics, Spark-to-Paper achieves 99.5% citation validity and 96.4% figure editability. A controlled ablation increases fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, while adversarial review achieves 74% precision. The full system uses 11.9M tokens, costs $8.1 per manuscript, and requires 3.2 hours on average. These results show that end-to-end research paper generation can be implemented as a lightweight, composable workflow inside existing coding assistants while keeping experimental evidence central to how claims are accepted, revised, or abandoned.","mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/62b32a4429a410b7f6b06710/pXhDS4-1KDERQMZFpgMo3.png"],"thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2608.11924.png","numComments":3,"upvoted":false,"submittedBy":{"_id":"62b32a4429a410b7f6b06710","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62b32a4429a410b7f6b06710/VzgvmnlYZWuifZTkIkCxy.jpeg","fullname":"Wenhao Wang","name":"WenhaoWang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":21,"isUserFollowing":false},"isAuthorParticipating":false},{"paper":{"id":"2607.24653","authors":[{"_id":"6a681e3e73f69d5af2bec64b","name":"Kimi Team","hidden":false},{"_id":"6a681e3e73f69d5af2bec64c","name":"Tongtong Bai","hidden":false},{"_id":"6a681e3e73f69d5af2bec64d","name":"Yifan Bai","hidden":false},{"_id":"6a681e3e73f69d5af2bec64e","name":"Yiping Bao","hidden":false},{"_id":"6a681e3e73f69d5af2bec64f","name":"M. C.","hidden":false},{"_id":"6a681e3e73f69d5af2bec650","name":"Jianfeng Cai","hidden":false},{"_id":"6a681e3e73f69d5af2bec651","name":"Xinyuan Cai","hidden":false},{"_id":"6a681e3e73f69d5af2bec652","name":"Peizhou Cao","hidden":false},{"_id":"6a681e3e73f69d5af2bec653","name":"Yuxuan Cao","hidden":false},{"_id":"6a681e3e73f69d5af2bec654","name":"Ziwei Chai","hidden":false},{"_id":"6a681e3e73f69d5af2bec655","name":"Y. Charles","hidden":false},{"_id":"6a681e3e73f69d5af2bec656","name":"H. S. Che","hidden":false},{"_id":"6a681e3e73f69d5af2bec657","name":"Guanduo Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec658","name":"Guangyu Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec659","name":"Guanzheng Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec65a","name":"Huarong Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec65b","name":"Jia Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec65c","name":"Jianlong Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec65d","name":"Jun Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec65e","name":"Kexin Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec65f","name":"Peng Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec660","name":"Ruijue Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec661","name":"Wentao Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec662","name":"Xin Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec663","name":"Yang Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec664","name":"Yanru Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec665","name":"Yifei Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec666","name":"Yingjiang Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec667","name":"Yuankun Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec668","name":"Yujie Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec669","name":"Yutian Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec66a","name":"Zhirong Chen","hidden":false},{"_id":"6a681e3e73f69d5af2bec66b","name":"Dazhi Cheng","hidden":false},{"_id":"6a681e3e73f69d5af2bec66c","name":"Yean Cheng","hidden":false},{"_id":"6a681e3e73f69d5af2bec66d","name":"Jialei Cui","hidden":false},{"_id":"6a681e3e73f69d5af2bec66e","name":"Jingbing Cui","hidden":false},{"_id":"6a681e3e73f69d5af2bec66f","name":"Anqi Dai","hidden":false},{"_id":"6a681e3e73f69d5af2bec670","user":{"_id":"66eeeb2ae65d94c88e9af620","avatarUrl":"/avatars/a25657d634878e9d53ada19feb38149a.svg","isPro":false,"fullname":"Jiaqi Deng","user":"MillanK","type":"user","name":"MillanK"},"name":"Jiaqi Deng","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.598Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec671","name":"Hao Ding","hidden":false},{"_id":"6a681e3e73f69d5af2bec672","name":"Rui Ding","hidden":false},{"_id":"6a681e3e73f69d5af2bec673","name":"Shaofeng Ding","hidden":false},{"_id":"6a681e3e73f69d5af2bec674","name":"Mengfan Dong","hidden":false},{"_id":"6a681e3e73f69d5af2bec675","name":"Mengnan Dong","hidden":false},{"_id":"6a681e3e73f69d5af2bec676","user":{"_id":"652965773a416e1f2173443b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652965773a416e1f2173443b/y9MB8YgHzbwCXAc4EI9T3.jpeg","isPro":false,"fullname":"Yuhao Dong","user":"THUdyh","type":"user","name":"THUdyh"},"name":"Yuhao Dong","status":"claimed_verified","statusLastChangedAt":"2026-07-29T08:45:04.322Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec677","name":"Yuxin Dong","hidden":false},{"_id":"6a681e3e73f69d5af2bec678","name":"Angang Du","hidden":false},{"_id":"6a681e3e73f69d5af2bec679","name":"Chenzhuang Du","hidden":false},{"_id":"6a681e3e73f69d5af2bec67a","name":"Dikang Du","hidden":false},{"_id":"6a681e3e73f69d5af2bec67b","user":{"_id":"65003e857804f04a163328d9","avatarUrl":"/avatars/fe32150aabfde8d283b38ccebcf6982e.svg","isPro":false,"fullname":"Jusen Du","user":"JusenK","type":"user","name":"JusenK"},"name":"Jusen Du","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.317Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec67c","name":"Yulun Du","hidden":false},{"_id":"6a681e3e73f69d5af2bec67d","name":"Yu Fan","hidden":false},{"_id":"6a681e3e73f69d5af2bec67e","name":"Jing Feng","hidden":false},{"_id":"6a681e3e73f69d5af2bec67f","name":"Qiulin Feng","hidden":false},{"_id":"6a681e3e73f69d5af2bec680","name":"Yichen Feng","hidden":false},{"_id":"6a681e3e73f69d5af2bec681","name":"Kelin Fu","hidden":false},{"_id":"6a681e3e73f69d5af2bec682","name":"Qiang Fu","hidden":false},{"_id":"6a681e3e73f69d5af2bec683","name":"Fuxuan Gao","hidden":false},{"_id":"6a681e3e73f69d5af2bec684","name":"Hongcheng Gao","hidden":false},{"_id":"6a681e3e73f69d5af2bec685","name":"Jingyue Gao","hidden":false},{"_id":"6a681e3e73f69d5af2bec686","name":"Tong Gao","hidden":false},{"_id":"6a681e3e73f69d5af2bec687","name":"Weijia Gao","hidden":false},{"_id":"6a681e3e73f69d5af2bec688","name":"Shangyi Geng","hidden":false},{"_id":"6a681e3e73f69d5af2bec689","name":"Jie Gong","hidden":false},{"_id":"6a681e3e73f69d5af2bec68a","name":"Linhu Gong","hidden":false},{"_id":"6a681e3e73f69d5af2bec68b","name":"Shengao Gong","hidden":false},{"_id":"6a681e3e73f69d5af2bec68c","name":"Xiaochen Gong","hidden":false},{"_id":"6a681e3e73f69d5af2bec68d","name":"Qizheng Gu","hidden":false},{"_id":"6a681e3e73f69d5af2bec68e","name":"Yicheng Gu","hidden":false},{"_id":"6a681e3e73f69d5af2bec68f","name":"Shuhao Guan","hidden":false},{"_id":"6a681e3e73f69d5af2bec690","name":"Haiqing Guo","hidden":false},{"_id":"6a681e3e73f69d5af2bec691","name":"Shiqi Guo","hidden":false},{"_id":"6a681e3e73f69d5af2bec692","name":"Xiang Guo","hidden":false},{"_id":"6a681e3e73f69d5af2bec693","name":"Zhengyan Guo","hidden":false},{"_id":"6a681e3e73f69d5af2bec694","name":"Beixi Hao","hidden":false},{"_id":"6a681e3e73f69d5af2bec695","name":"Wenxin Hao","hidden":false},{"_id":"6a681e3e73f69d5af2bec696","name":"Xiaoru Hao","hidden":false},{"_id":"6a681e3e73f69d5af2bec697","name":"Dailan He","hidden":false},{"_id":"6a681e3e73f69d5af2bec698","name":"Haotian He","hidden":false},{"_id":"6a681e3e73f69d5af2bec699","user":{"_id":"6753c09d95f2fead90a3750a","avatarUrl":"/avatars/6d7db74a47cbb4206dd3796632f32936.svg","isPro":false,"fullname":"helehan(SII)","user":"helehan","type":"user","name":"helehan"},"name":"Lehan He","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:58:06.831Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec69a","name":"Qi He","hidden":false},{"_id":"6a681e3e73f69d5af2bec69b","name":"Weiran He","hidden":false},{"_id":"6a681e3e73f69d5af2bec69c","name":"Xinran He","hidden":false},{"_id":"6a681e3e73f69d5af2bec69d","name":"Xinyi He","hidden":false},{"_id":"6a681e3e73f69d5af2bec69e","name":"Yibo He","hidden":false},{"_id":"6a681e3e73f69d5af2bec69f","name":"Yunjia He","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a0","name":"Chao Hong","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a1","name":"Tiange Hong","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a2","name":"Hao Hu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a3","name":"Jiaxi Hu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a4","name":"Ruikun Hu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a5","name":"Weiming Hu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a6","name":"Yangyang Hu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a7","name":"Zhenxing Hu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a8","name":"Liang Hua","hidden":false},{"_id":"6a681e3e73f69d5af2bec6a9","name":"Jinbin Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6aa","name":"Ke Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ab","name":"Ruiyuan Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ac","name":"Siying Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ad","name":"Weixiao Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ae","name":"Yan Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6af","name":"Zhengjie Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b0","name":"Zhiqi Huang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b1","name":"Yulong Hui","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b2","name":"Chaobo Jia","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b3","name":"Yutong Jiang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b4","name":"Zhejun Jiang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b5","name":"Zuoyou Jiang","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b6","name":"Wenyi Jin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b7","name":"Xinyi Jin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b8","name":"Yu Jing","hidden":false},{"_id":"6a681e3e73f69d5af2bec6b9","name":"Huanjun Kong","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ba","name":"Guokun Lai","hidden":false},{"_id":"6a681e3e73f69d5af2bec6bb","name":"Aidi Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6bc","name":"Cheng Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6bd","name":"Chengyuan Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6be","name":"Cong Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6bf","name":"Fang Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c0","name":"Guanyu Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c1","name":"Haoyang Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c2","name":"Jia Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c3","name":"Junxiong Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c4","name":"Lei Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c5","name":"Letian Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c6","name":"Lincan Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c7","name":"Weihong Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c8","name":"Wentao Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6c9","name":"Xintong Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ca","name":"Yang Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6cb","name":"Yishen Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6cc","name":"Yiwei Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6cd","name":"Yuxiao Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ce","name":"Zhaowei Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6cf","name":"Zhaoxi Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d0","name":"Zheming Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d1","name":"Zhengxiao Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d2","name":"Zhiyuan Li","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d3","name":"Jiawei Lin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d4","name":"Xiaohan Lin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d5","name":"Yibo Lin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d6","name":"Zichao Lin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d7","name":"Ziyan Lin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d8","name":"Bill Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6d9","name":"Boxiao Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6da","name":"Chuan Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6db","name":"Liang Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6dc","name":"Shaowei Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6dd","user":{"_id":"654ce87af0b05673196a9f45","avatarUrl":"/avatars/7b9c854eb98e487e3057479b1c7860ac.svg","isPro":false,"fullname":"Shudong Liu","user":"Sudanl","type":"user","name":"Sudanl"},"name":"Shudong Liu","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.624Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec6de","name":"Shuran Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6df","name":"Tianwei Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e0","name":"Weizhou Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e1","name":"Yangyang Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e2","name":"Yanming Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e3","name":"Yibo Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e4","name":"Yipeng Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e5","name":"Zhengying Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e6","name":"Zhiheng Liu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e7","name":"Enzhe Lu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e8","name":"Haoyu Lu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6e9","name":"Linqiang Lu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ea","name":"Tingzhan Lu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6eb","name":"Zhiyuan Lu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ec","name":"Aotian Luo","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ed","name":"G. Luo","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ee","user":{"_id":"642da1cd99f3110ac27caca5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642da1cd99f3110ac27caca5/C1QJY3R_ZdaeANG1y8iW7.jpeg","isPro":false,"fullname":"junyu","user":"luojunyu","type":"user","name":"luojunyu"},"name":"Junyu Luo","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:58:09.894Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ef","user":{"_id":"624dae2cb58d1313f8ba9a2d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/624dae2cb58d1313f8ba9a2d/cFar_YWV8qtWUjOZeNKT0.png","isPro":false,"fullname":"Yifan Luo","user":"Luobots","type":"user","name":"Luobots"},"name":"Yifan Luo","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.605Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f0","name":"B. Lyu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f1","name":"Wenzhou Lyu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f2","user":{"_id":"64af8d4b2cda6a37a4927d72","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/o37G6VkKSyP4N9xktFErV.png","isPro":false,"fullname":"Shaoguang Mao","user":"dawnmsg","type":"user","name":"dawnmsg"},"name":"Shaoguang Mao","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.611Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f3","name":"Yuan Mei","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f4","name":"Xin Men","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f5","name":"Minqing Ni","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f6","name":"Yixuan Niu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f7","name":"Siyuan Pan","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f8","name":"Shujun Peng","hidden":false},{"_id":"6a681e3e73f69d5af2bec6f9","name":"Zhangyang Qi","hidden":false},{"_id":"6a681e3e73f69d5af2bec6fa","name":"Ruoyu Qin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6fb","name":"ZeChao Qin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6fc","name":"Zeyu Qin","hidden":false},{"_id":"6a681e3e73f69d5af2bec6fd","name":"Haiquan Qiu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6fe","name":"Jianxin Qiu","hidden":false},{"_id":"6a681e3e73f69d5af2bec6ff","name":"Jiezhong Qiu","hidden":false},{"_id":"6a681e3e73f69d5af2bec700","name":"Bowen Qu","hidden":false},{"_id":"6a681e3e73f69d5af2bec701","name":"Yuhao Qu","hidden":false},{"_id":"6a681e3e73f69d5af2bec702","name":"Zeyu Shang","hidden":false},{"_id":"6a681e3e73f69d5af2bec703","name":"Youbo Shao","hidden":false},{"_id":"6a681e3e73f69d5af2bec704","name":"Han Shen","hidden":false},{"_id":"6a681e3e73f69d5af2bec705","name":"Jincheng Shi","hidden":false},{"_id":"6a681e3e73f69d5af2bec706","name":"Juanfeng Shi","hidden":false},{"_id":"6a681e3e73f69d5af2bec707","name":"Lidong Shi","hidden":false},{"_id":"6a681e3e73f69d5af2bec708","name":"Shengyuan Shi","hidden":false},{"_id":"6a681e3e73f69d5af2bec709","name":"Wingchun Siu","hidden":false},{"_id":"6a681e3e73f69d5af2bec70a","name":"Pengwei Song","hidden":false},{"_id":"6a681e3e73f69d5af2bec70b","name":"Xiaoxi Song","hidden":false},{"_id":"6a681e3e73f69d5af2bec70c","name":"Jianlin Su","hidden":false},{"_id":"6a681e3e73f69d5af2bec70d","name":"Yunfeng Su","hidden":false},{"_id":"6a681e3e73f69d5af2bec70e","user":{"_id":"64264095ba51f8a2136946a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64264095ba51f8a2136946a0/FR33boVpkDXcrvGMBmprF.jpeg","isPro":false,"fullname":"Zhaochen Su","user":"Warrieryes","type":"user","name":"Warrieryes"},"name":"Zhaochen Su","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.618Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec70f","name":"Lin Sui","hidden":false},{"_id":"6a681e3e73f69d5af2bec710","name":"Jingsong Sun","hidden":false},{"_id":"6a681e3e73f69d5af2bec711","name":"Junyao Sun","hidden":false},{"_id":"6a681e3e73f69d5af2bec712","name":"Shaoning Sun","hidden":false},{"_id":"6a681e3e73f69d5af2bec713","name":"Shuzhe Sun","hidden":false},{"_id":"6a681e3e73f69d5af2bec714","name":"Tongyu Sun","hidden":false},{"_id":"6a681e3e73f69d5af2bec715","name":"Yujun Sun","hidden":false},{"_id":"6a681e3e73f69d5af2bec716","name":"Yunpeng Tai","hidden":false},{"_id":"6a681e3e73f69d5af2bec717","name":"Chuning Tang","hidden":false},{"_id":"6a681e3e73f69d5af2bec718","name":"Heyi Tang","hidden":false},{"_id":"6a681e3e73f69d5af2bec719","name":"Sirui Tang","hidden":false},{"_id":"6a681e3e73f69d5af2bec71a","name":"Zecheng Tang","hidden":false},{"_id":"6a681e3e73f69d5af2bec71b","name":"Chaoran Tian","hidden":false},{"_id":"6a681e3e73f69d5af2bec71c","name":"Rongpeng Tian","hidden":false},{"_id":"6a681e3e73f69d5af2bec71d","name":"Yu Tian","hidden":false},{"_id":"6a681e3e73f69d5af2bec71e","name":"Wei Tu","hidden":false},{"_id":"6a681e3e73f69d5af2bec71f","name":"Chensi Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec720","name":"Chuang Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec721","name":"Chunjie Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec722","name":"Dinglu Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec723","name":"Feng Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec724","name":"Hailong Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec725","name":"Haiming Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec726","name":"Hao Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec727","name":"Hao Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec728","name":"Huaqing Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec729","name":"Hui Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec72a","name":"Jiayi Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec72b","name":"Jinglong Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec72c","name":"Jinhong Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec72d","name":"Jiuzheng Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec72e","name":"Linian Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec72f","user":{"_id":"66968099c952e09a4cb29f78","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66968099c952e09a4cb29f78/n90NI2R3E9_RqCyMjDCQF.webp","isPro":false,"fullname":"Wang","user":"Steven-Shaobo","type":"user","name":"Steven-Shaobo"},"name":"Shaobo Wang","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:58:12.154Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec730","name":"Shenzhi Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec731","name":"Shuyi Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec732","name":"Si Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec733","name":"Siyuan Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec734","name":"Tianfu Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec735","name":"Wenjue Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec736","name":"Xingran Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec737","name":"Xinmei Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec738","user":{"_id":"67b327cdd4665a0448eef7d5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67b327cdd4665a0448eef7d5/_B5Z9MCa_qiFrDj1axKlz.png","isPro":true,"fullname":"Xinyuan Wang","user":"xywang626","type":"user","name":"xywang626"},"name":"Xinyuan Wang","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.309Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec739","name":"Xusheng Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec73a","name":"Yalin Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec73b","name":"Yangkun Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec73c","name":"Yao Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec73d","name":"Yaoyu Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec73e","name":"Yejie Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec73f","name":"Yiqin Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec740","name":"Yucheng Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec741","name":"Yuzhi Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec742","name":"Zhaoji Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec743","name":"Zhaowei Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec744","name":"Zhengtao Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec745","name":"Zhenhao Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec746","name":"Zhongsheng Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec747","name":"Zifan Wang","hidden":false},{"_id":"6a681e3e73f69d5af2bec748","user":{"_id":"635ddec594e5b275ca7941e8","avatarUrl":"/avatars/28ebfaee74d31e1de020a3ae735a4c1b.svg","isPro":false,"fullname":"Chu Wei","user":"courage17340","type":"user","name":"courage17340"},"name":"Chu Wei","status":"claimed_verified","statusLastChangedAt":"2026-08-04T08:45:04.597Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec749","name":"Ming Wei","hidden":false},{"_id":"6a681e3e73f69d5af2bec74a","name":"Shouxin Wei","hidden":false},{"_id":"6a681e3e73f69d5af2bec74b","name":"Zichen Wen","hidden":false},{"_id":"6a681e3e73f69d5af2bec74c","name":"Fan Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec74d","name":"Haoning Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec74e","name":"Rucong Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec74f","name":"Wenhao Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec750","name":"Xiaoxue Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec751","name":"Yingcong Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec752","name":"Yongqi Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec753","name":"Yuxin Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec754","name":"Zijian Wu","hidden":false},{"_id":"6a681e3e73f69d5af2bec755","name":"Xinglang Xian","hidden":false},{"_id":"6a681e3e73f69d5af2bec756","name":"Chenxuan Xiang","hidden":false},{"_id":"6a681e3e73f69d5af2bec757","name":"Yuye Xiang","hidden":false},{"_id":"6a681e3e73f69d5af2bec758","name":"Bocheng Xiao","hidden":false},{"_id":"6a681e3e73f69d5af2bec759","name":"Chenjun Xiao","hidden":false},{"_id":"6a681e3e73f69d5af2bec75a","name":"Xin Xiao","hidden":false},{"_id":"6a681e3e73f69d5af2bec75b","name":"Jin Xie","hidden":false},{"_id":"6a681e3e73f69d5af2bec75c","name":"Xiaotong Xie","hidden":false},{"_id":"6a681e3e73f69d5af2bec75d","name":"Yifeng Xie","hidden":false},{"_id":"6a681e3e73f69d5af2bec75e","name":"Zhe Xie","hidden":false},{"_id":"6a681e3e73f69d5af2bec75f","name":"Bowei Xing","hidden":false},{"_id":"6a681e3e73f69d5af2bec760","name":"Yiming Xiong","hidden":false},{"_id":"6a681e3e73f69d5af2bec761","name":"Baosheng Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec762","name":"Boyu Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec763","name":"Jiale Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec764","name":"Jianfan Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec765","name":"Jing Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec766","name":"Jinjing Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec767","name":"L. H. Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec768","name":"Qingtao Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec769","name":"Shuyao Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec76a","name":"Suting Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec76b","name":"Tiantian Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec76c","name":"Tianxiang Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec76d","name":"Weixin Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec76e","name":"Xinran Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec76f","name":"Yangchuan Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec770","name":"Ye Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec771","name":"Yueni Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec772","name":"Ziyao Xu","hidden":false},{"_id":"6a681e3e73f69d5af2bec773","name":"Haonan Xue","hidden":false},{"_id":"6a681e3e73f69d5af2bec774","name":"Junjie Yan","hidden":false},{"_id":"6a681e3e73f69d5af2bec775","name":"Yaoyao Yan","hidden":false},{"_id":"6a681e3e73f69d5af2bec776","name":"Fan Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec777","name":"Guangyao Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec778","name":"Hao Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec779","name":"Junwei Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec77a","name":"Ruoyu Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec77b","name":"Wenjie Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec77c","name":"Xiaofei Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec77d","name":"Xinyu Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec77e","name":"Yi Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec77f","name":"Yiling Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec780","name":"Ying Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec781","name":"Yuchen Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec782","name":"Zhen Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec783","name":"Zhilin Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec784","name":"Zian Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec785","name":"Zuhao Yang","hidden":false},{"_id":"6a681e3e73f69d5af2bec786","name":"Haotian Yao","hidden":false},{"_id":"6a681e3e73f69d5af2bec787","name":"Dan Ye","hidden":false},{"_id":"6a681e3e73f69d5af2bec788","name":"Haoran Ye","hidden":false},{"_id":"6a681e3e73f69d5af2bec789","name":"Wenjie Ye","hidden":false},{"_id":"6a681e3e73f69d5af2bec78a","name":"Zhanbo Ye","hidden":false},{"_id":"6a681e3e73f69d5af2bec78b","name":"Bohong Yin","hidden":false},{"_id":"6a681e3e73f69d5af2bec78c","name":"Haoxiang Yin","hidden":false},{"_id":"6a681e3e73f69d5af2bec78d","name":"Xietong Yin","hidden":false},{"_id":"6a681e3e73f69d5af2bec78e","name":"Chengzhen Yu","hidden":false},{"_id":"6a681e3e73f69d5af2bec78f","name":"Haozhen Yu","hidden":false},{"_id":"6a681e3e73f69d5af2bec790","name":"Longhui Yu","hidden":false},{"_id":"6a681e3e73f69d5af2bec791","name":"Shengnan Yu","hidden":false},{"_id":"6a681e3e73f69d5af2bec792","name":"Shuying Yu","hidden":false},{"_id":"6a681e3e73f69d5af2bec793","name":"Tianxiang Yu","hidden":false},{"_id":"6a681e3e73f69d5af2bec794","name":"Enming Yuan","hidden":false},{"_id":"6a681e3e73f69d5af2bec795","name":"Mengjie Yuan","hidden":false},{"_id":"6a681e3e73f69d5af2bec796","name":"Tongtian Yue","hidden":false},{"_id":"6a681e3e73f69d5af2bec797","name":"Wei Yue","hidden":false},{"_id":"6a681e3e73f69d5af2bec798","name":"Yang Yue","hidden":false},{"_id":"6a681e3e73f69d5af2bec799","name":"Dunyuan Zha","hidden":false},{"_id":"6a681e3e73f69d5af2bec79a","name":"Haobing Zhan","hidden":false},{"_id":"6a681e3e73f69d5af2bec79b","name":"B. H. Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec79c","name":"Dehao Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec79d","name":"Fei Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec79e","name":"Hao Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec79f","name":"Haoyuan Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a0","name":"Huanyu Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a1","name":"Jiapei Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a2","name":"Jiaxuan Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a3","name":"Jin Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a4","name":"Kaiyi Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a5","name":"Miaozhen Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a6","name":"Puqi Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a7","name":"Qinglei Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a8","name":"Rong Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7a9","name":"Rui Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7aa","name":"Shaoshuai Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7ab","name":"Shiyi Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7ac","name":"Xiaobin Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7ad","user":{"_id":"67f33b43ccb05db5f0190cf6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/c-UeSyHfO7QchwNXW9mgu.png","isPro":false,"fullname":"ZhangXiaoyun","user":"DadaCloud01","type":"user","name":"DadaCloud01"},"name":"Xiaoyun Zhang","status":"claimed_verified","statusLastChangedAt":"2026-07-29T08:45:04.331Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec7ae","name":"Y. Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7af","name":"Yangkun Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b0","name":"Ye Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b1","name":"Yichi Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b2","name":"Yikun Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b3","name":"Yizhi Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b4","name":"Yongting Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b5","user":{"_id":"638f5839f6de4b9e7e1627fb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638f5839f6de4b9e7e1627fb/6QGkrqRag6-GnH9k60Oil.jpeg","isPro":false,"fullname":"Yu Zhang","user":"yzhangcs","type":"user","name":"yzhangcs"},"name":"Yu Zhang","status":"claimed_verified","statusLastChangedAt":"2026-07-28T16:45:04.639Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b6","name":"Yutao Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b7","name":"Yutong Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b8","name":"Zheng Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7b9","name":"Zijing Zhang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7ba","name":"Bin Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7bb","name":"Chenguang Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7bc","name":"Feifan Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7bd","name":"Jinglun Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7be","name":"Jinxiang Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7bf","name":"Shuai Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c0","name":"Wenshuo Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c1","name":"Xiangyu Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c2","name":"Xuanle Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c3","name":"Yikai Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c4","name":"Zijia Zhao","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c5","name":"Haozhi Zheng","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c6","name":"Huabin Zheng","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c7","name":"Ruihan Zheng","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c8","name":"Shaojie Zheng","hidden":false},{"_id":"6a681e3e73f69d5af2bec7c9","name":"Tengyang Zheng","hidden":false},{"_id":"6a681e3e73f69d5af2bec7ca","name":"Haofeng Zhong","hidden":false},{"_id":"6a681e3e73f69d5af2bec7cb","name":"Lei Zhong","hidden":false},{"_id":"6a681e3e73f69d5af2bec7cc","user":{"_id":"62b6d20416ff90e6198301b6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1656148456743-noauth.png","isPro":false,"fullname":"Longguang Zhong","user":"GGLS","type":"user","name":"GGLS"},"name":"Longguang Zhong","status":"claimed_verified","statusLastChangedAt":"2026-07-28T16:45:04.647Z","hidden":false},{"_id":"6a681e3e73f69d5af2bec7cd","name":"M. Zhou","hidden":false},{"_id":"6a681e3e73f69d5af2bec7ce","name":"Qiankang Zhou","hidden":false},{"_id":"6a681e3e73f69d5af2bec7cf","name":"Runjie Zhou","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d0","name":"Ruozhang Zhou","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d1","name":"Xinyu Zhou","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d2","name":"Yiqiao Zhou","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d3","name":"Zaida Zhou","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d4","name":"Jinguo Zhu","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d5","name":"Liya Zhu","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d6","name":"Xinhao Zhu","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d7","name":"Yangjunfeng Zhu","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d8","name":"Yuxuan Zhu","hidden":false},{"_id":"6a681e3e73f69d5af2bec7d9","name":"Zhen Zhu","hidden":false},{"_id":"6a681e3e73f69d5af2bec7da","name":"Chen Zhuang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7db","name":"Weiyu Zhuang","hidden":false},{"_id":"6a681e3e73f69d5af2bec7dc","name":"Xinxing Zu","hidden":false}],"publishedAt":"2026-07-27T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"Kimi K3: Open Frontier Intelligence","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.","upvotes":493,"discussionId":"6a681e3e73f69d5af2bec7dd","projectPage":"https://www.kimi.com/blog/kimi-k3","githubRepo":"https://github.com/MoonshotAI/Kimi-K3","githubRepoAddedBy":"admin","ai_summary":"Kimi K3 is a large-scale mixture-of-experts model with native vision and long-context capabilities that improves scaling efficiency and achieves strong performance across coding, reasoning, and agentic tasks.","ai_keywords":["Mixture-of-Experts","Kimi Delta Attention","Attention Residuals","Stable LatentMoE","expert-parallel training","reinforcement learning","long-horizon execution","compositional generalization","vision capabilities"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8486,"organization":{"_id":"6425a114812813f8f4a9b02c","name":"moonshotai","fullname":"Moonshot AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/641c1e77c3983aa9490f8121/X1yT2rsaIbR9cdYGEVu0X.jpeg"}},"publishedAt":"2026-07-26T20:00:00.000Z","title":"Kimi K3: Open Frontier Intelligence","summary":"We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2607.24653.png","numComments":10,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"organization":{"_id":"6425a114812813f8f4a9b02c","name":"moonshotai","fullname":"Moonshot AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/641c1e77c3983aa9490f8121/X1yT2rsaIbR9cdYGEVu0X.jpeg"},"isAuthorParticipating":false},{"paper":{"id":"2403.13372","authors":[{"_id":"65fbae8adfd0ff86aa7a217a","user":{"_id":"642fef28a043f0ac7defa8a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/RwOEkuj3fOnOA54tGR7Ea.png","isPro":false,"fullname":"Yaowei Zheng","user":"hiyouga","type":"user","name":"hiyouga"},"name":"Yaowei Zheng","status":"admin_assigned","statusLastChangedAt":"2024-03-21T09:11:34.746Z","hidden":false},{"_id":"65fbae8adfd0ff86aa7a217b","name":"Richong Zhang","hidden":false},{"_id":"65fbae8adfd0ff86aa7a217c","user":{"_id":"6524d66966ebe0519873c4c5","avatarUrl":"/avatars/18a52cf9cf90f9ff309284a1c9cae6ac.svg","isPro":false,"fullname":"Junhao Zhang","user":"OnlyAR","type":"user","name":"OnlyAR"},"name":"Junhao Zhang","status":"admin_assigned","statusLastChangedAt":"2024-03-21T09:33:46.449Z","hidden":false},{"_id":"65fbae8adfd0ff86aa7a217d","user":{"_id":"65fc00206529e3fcc230b820","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65fc00206529e3fcc230b820/IeZjMvHYZeZn8-igQ45R8.jpeg","isPro":false,"fullname":"Yanhan Ye","user":"CoolColoury","type":"user","name":"CoolColoury"},"name":"Yanhan Ye","status":"claimed_verified","statusLastChangedAt":"2024-03-21T10:02:39.082Z","hidden":false},{"_id":"65fbae8adfd0ff86aa7a217e","name":"Zheyan Luo","hidden":false}],"publishedAt":"2024-03-20T08:08:54.000Z","submittedOnDailyAt":"2024-03-21T00:00:00.000Z","title":"LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models","submittedOnDailyBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","isPro":false,"fullname":"AK","user":"akhaliq","type":"user","name":"akhaliq"},"summary":"Efficient fine-tuning is vital for adapting large language models (LLMs) to\ndownstream tasks. However, it requires non-trivial efforts to implement these\nmethods on different models. We present LlamaFactory, a unified framework that\nintegrates a suite of cutting-edge efficient training methods. It allows users\nto flexibly customize the fine-tuning of 100+ LLMs without the need for coding\nthrough the built-in web UI LlamaBoard. We empirically validate the efficiency\nand effectiveness of our framework on language modeling and text generation\ntasks. It has been released at https://github.com/hiyouga/LLaMA-Factory and\nalready received over 13,000 stars and 1,600 forks.","upvotes":187,"discussionId":"65fbae8bdfd0ff86aa7a219d","projectPage":"https://huggingface.co/spaces/hiyouga/LLaMA-Board","githubRepo":"https://github.com/hiyouga/LLaMA-Factory","githubRepoAddedBy":"user","ai_summary":"LlamaFactory is a unified framework enabling efficient fine-tuning of large language models across various tasks using a web-based user interface.","ai_keywords":["efficient fine-tuning","large language models","LLaMA","LlamaFactory","LlamaBoard","language modeling","text generation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":74158},"publishedAt":"2024-03-20T04:08:54.000Z","title":"LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models","summary":"Efficient fine-tuning is vital for adapting large language models (LLMs) to\ndownstream tasks. However, it requires non-trivial efforts to implement these\nmethods on different models. We present LlamaFactory, a unified framework that\nintegrates a suite of cutting-edge efficient training methods. It allows users\nto flexibly customize the fine-tuning of 100+ LLMs without the need for coding\nthrough the built-in web UI LlamaBoard. We empirically validate the efficiency\nand effectiveness of our framework on language modeling and text generation\ntasks. It has been released at https://github.com/hiyouga/LLaMA-Factory and\nalready received over 13,000 stars and 1,600 forks.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2403.13372.png","numComments":6,"upvoted":false,"submittedBy":{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","fullname":"AK","name":"akhaliq","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":9966,"isUserFollowing":false},"isAuthorParticipating":false},{"paper":{"id":"2508.16279","authors":[{"_id":"68abe18a86b21a0e2e358adf","name":"Dawei Gao","hidden":false},{"_id":"68abe18a86b21a0e2e358ae0","name":"Zitao Li","hidden":false},{"_id":"68abe18a86b21a0e2e358ae1","name":"Yuexiang Xie","hidden":false},{"_id":"68abe18a86b21a0e2e358ae2","user":{"_id":"62734fc860bc37aa9c326f50","avatarUrl":"/avatars/2238451f16559c7a6d77f3dcd6fcd1f3.svg","isPro":false,"fullname":"weirui kuang","user":"rayrayray","type":"user","name":"rayrayray"},"name":"Weirui Kuang","status":"admin_assigned","statusLastChangedAt":"2025-08-25T20:34:06.130Z","hidden":false},{"_id":"68abe18a86b21a0e2e358ae3","name":"Liuyi Yao","hidden":false},{"_id":"68abe18a86b21a0e2e358ae4","user":{"_id":"64d0572ec0c627dfa7e98e03","avatarUrl":"/avatars/7172533bcc26ae9222133e7efdf5e0b2.svg","isPro":false,"fullname":"qbc","user":"qbc","type":"user","name":"qbc"},"name":"Bingchen Qian","status":"claimed_verified","statusLastChangedAt":"2025-09-15T15:09:33.146Z","hidden":false},{"_id":"68abe18a86b21a0e2e358ae5","name":"Zhijian Ma","hidden":false},{"_id":"68abe18a86b21a0e2e358ae6","name":"Yue Cui","hidden":false},{"_id":"68abe18a86b21a0e2e358ae7","name":"Haohao Luo","hidden":false},{"_id":"68abe18a86b21a0e2e358ae8","user":{"_id":"64cc7dd50d0e2410eeaad6d3","avatarUrl":"/avatars/9a31e435574f8ca2a926c39019aae8ce.svg","isPro":false,"fullname":"Shen Li","user":"Listen0425","type":"user","name":"Listen0425"},"name":"Shen Li","status":"claimed_verified","statusLastChangedAt":"2025-08-26T09:50:34.184Z","hidden":false},{"_id":"68abe18a86b21a0e2e358ae9","name":"Lu Yi","hidden":false},{"_id":"68abe18a86b21a0e2e358aea","name":"Yi Yu","hidden":false},{"_id":"68abe18a86b21a0e2e358aeb","name":"Shiqi He","hidden":false},{"_id":"68abe18a86b21a0e2e358aec","name":"Zhiling Luo","hidden":false},{"_id":"68abe18a86b21a0e2e358aed","name":"Wenmeng Zhou","hidden":false},{"_id":"68abe18a86b21a0e2e358aee","name":"Zhicheng Zhang","hidden":false},{"_id":"68abe18a86b21a0e2e358aef","name":"Xuguang He","hidden":false},{"_id":"68abe18a86b21a0e2e358af0","name":"Ziqian Chen","hidden":false},{"_id":"68abe18a86b21a0e2e358af1","name":"Weikai Liao","hidden":false},{"_id":"68abe18a86b21a0e2e358af2","name":"Farruh Isakulovich Kushnazarov","hidden":false},{"_id":"68abe18a86b21a0e2e358af3","name":"Yaliang Li","hidden":false},{"_id":"68abe18a86b21a0e2e358af4","name":"Bolin Ding","hidden":false},{"_id":"68abe18a86b21a0e2e358af5","name":"Jingren Zhou","hidden":false}],"publishedAt":"2025-08-22T10:35:56.000Z","submittedOnDailyAt":"2025-08-25T00:00:00.000Z","title":"AgentScope 1.0: A Developer-Centric Framework for Building Agentic\n Applications","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Driven by rapid advancements of Large Language Models (LLMs), agents are\nempowered to combine intrinsic knowledge with dynamic tool use, greatly\nenhancing their capacity to address real-world tasks. In line with such an\nevolution, AgentScope introduces major improvements in a new version (1.0),\ntowards comprehensively supporting flexible and efficient tool-based\nagent-environment interactions for building agentic applications. Specifically,\nwe abstract foundational components essential for agentic applications and\nprovide unified interfaces and extensible modules, enabling developers to\neasily leverage the latest progress, such as new models and MCPs. Furthermore,\nwe ground agent behaviors in the ReAct paradigm and offer advanced agent-level\ninfrastructure based on a systematic asynchronous design, which enriches both\nhuman-agent and agent-agent interaction patterns while improving execution\nefficiency. Building on this foundation, we integrate several built-in agents\ntailored to specific practical scenarios. AgentScope also includes robust\nengineering support for developer-friendly experiences. We provide a scalable\nevaluation module with a visual studio interface, making the development of\nlong-trajectory agentic applications more manageable and easier to trace. In\naddition, AgentScope offers a runtime sandbox to ensure safe agent execution\nand facilitates rapid deployment in production environments. With these\nenhancements, AgentScope provides a practical foundation for building scalable,\nadaptive, and effective agentic applications.","upvotes":68,"discussionId":"68abe18b86b21a0e2e358af6","githubRepo":"https://github.com/agentscope-ai/agentscope","githubRepoAddedBy":"user","ai_summary":"AgentScope enhances agentic applications by providing flexible tool-based interactions, unified interfaces, and advanced infrastructure based on the ReAct paradigm, supporting efficient and safe development and deployment.","ai_keywords":["Large Language Models","LLMs","AgentScope","tool-based agent-environment interactions","ReAct paradigm","asynchronous design","human-agent interactions","agent-agent interactions","built-in agents","scalable evaluation module","visual studio interface","runtime sandbox"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":28994},"publishedAt":"2025-08-22T06:35:56.000Z","title":"AgentScope 1.0: A Developer-Centric Framework for Building Agentic\n Applications","summary":"Driven by rapid advancements of Large Language Models (LLMs), agents are\nempowered to combine intrinsic knowledge with dynamic tool use, greatly\nenhancing their capacity to address real-world tasks. In line with such an\nevolution, AgentScope introduces major improvements in a new version (1.0),\ntowards comprehensively supporting flexible and efficient tool-based\nagent-environment interactions for building agentic applications. Specifically,\nwe abstract foundational components essential for agentic applications and\nprovide unified interfaces and extensible modules, enabling developers to\neasily leverage the latest progress, such as new models and MCPs. Furthermore,\nwe ground agent behaviors in the ReAct paradigm and offer advanced agent-level\ninfrastructure based on a systematic asynchronous design, which enriches both\nhuman-agent and agent-agent interaction patterns while improving execution\nefficiency. Building on this foundation, we integrate several built-in agents\ntailored to specific practical scenarios. AgentScope also includes robust\nengineering support for developer-friendly experiences. We provide a scalable\nevaluation module with a visual studio interface, making the development of\nlong-trajectory agentic applications more manageable and easier to trace. In\naddition, AgentScope offers a runtime sandbox to ensure safe agent execution\nand facilitates rapid deployment in production environments. With these\nenhancements, AgentScope provides a practical foundation for building scalable,\nadaptive, and effective agentic applications.","thumbnail":"https://cdn-thumbnails.huggingface.co/social-thumbnails/papers/2508.16279.png","numComments":4,"upvoted":false,"submittedBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":359,"isUserFollowing":false},"isAuthorParticipating":false},{"paper":{"id":"2407.17789","authors":[{"_id":"66a31f39a74b20aad777e7e7","user":{"_id":"64101b4ca92fedb0e8559b91","avatarUrl":"/avatars/c7453a5641a94d04a75de9b7d4272bc2.svg","isPro":false,"fullname":"Xuchen Pan","user":"panxuchen","type":"user","name":"panxuchen"},"name":"Xuchen Pan","status":"admin_assigned","statusLastChangedAt":"2024-07-26T08:30:22.586Z","hidden":false},{"_id":"66a31f39a74b20aad777e7e8","user":{"_id":"642e79a83a486fec83426a49","avatarUrl":"/avatars/a003d27177cee6234348008263933cfd.svg","isPro":false,"fullname":"dawei gao","user":"DavdGao","type":"user","name":"DavdGao"},"name":"Dawei Gao","status":"admin_assigned","statusLastChangedAt":"2024-07-26T08:30:28.568Z","hidden":false},{"_id":"66a31f39a74b20aad777e7e9","name":"Yuexiang Xie","hidden":false},{"_id":"66a31f39a74b20aad777e7ea","name":"Zhewei Wei","hidden":false},{"_id":"66a31f39a74b20aad777e7eb","name":"Yaliang Li","hidden":false},{"_id":"66a31f39a74b20aad777e7ec","name":"Bolin Ding","hidden":false},{"_id":"66a31f39a74b20aad777e7ed","user":{"_id":"64b8c89052b7353d8c6a1013","avatarUrl":"/avatars/cd59fffe81f6b07b4519540b8ff3d95f.svg","isPro":false,"fullname":"Ji-Rong Wen","user":"jrwen","type":"user","name":"jrwen"},"name":"Ji-Rong Wen","status":"admin_assigned","statusLastChangedAt":"2024-07-26T08:32:01.623Z","hidden":false},{"_id":"66a31f39a74b20aad777e7ee",&q