-
Syed Waseem Haider, Ph.D.
IZ Analytics • 4K followers
Many students and young professionals reach out for career advice and are curious about future-proof skills. I often find myself recommending a strong foundation in mathematics, algorithms, physics, electrical engineering, signal processing, and probability theory, even as AI’s appeal commands the spotlight. We exist in a world governed by omnipresent forces, laws, constants, and constraints that define and bound our systems—whether they are biological, computational, economic, social, or physical. AI is no exception to this reality. Here’s what renowned researchers and leaders in AI think: https://lnkd.in/eJXW5c9A
-
Taha Ahmad
Cortessa Co • 915 followers
Excited to share that my latest research paper is now live on arXiv. 🚀 We benchmarked Similarity Network Fusion (SNF) against single-omic, early-integration, and late-integration baselines on TCGA-BRCA, integrating: • RNA-seq • 450k DNA methylation (ComBat batch-corrected) • GISTIC2 copy number data The pipeline includes spectral clustering, survival modeling (Ridge Cox), covariate-adjusted analysis, and robustness checks across multiple sensitivity settings. Across subtype recovery, cluster stability, and survival prediction, SNF consistently outperformed baseline approaches. This project reflects my broader interest in multi-omics integration, network methods, and computational oncology. Preprint: https://lnkd.in/d6f6Qr8N Code and pipeline: https://lnkd.in/d7DP3cxD
5 Comments
-
Shiva Krishna Reddy Malay
ServiceNow • 994 followers
Thrilled to share that our paper, “𝘈𝘙𝘔: 𝘋𝘪𝘴𝘤𝘰𝘷𝘦𝘳𝘪𝘯𝘨 𝘈𝘨𝘦𝘯𝘵𝘪𝘤 𝘙𝘦𝘢𝘴𝘰𝘯𝘪𝘯𝘨 𝘔𝘰𝘥𝘶𝘭𝘦𝘴 𝘧𝘰𝘳 𝘔𝘢𝘵𝘩𝘦𝘮𝘢𝘵𝘪𝘤𝘢𝘭 𝘗𝘳𝘰𝘣𝘭𝘦𝘮-𝘚𝘰𝘭𝘷𝘪𝘯𝘨,” was presented this past weekend at the NeurIPS MATH-AI Workshop! 🌟 𝗠𝗼𝘁𝗶𝘃𝗮𝘁𝗶𝗼𝗻: Automated search for multi-agentic architectures is an emerging frontier for inference time scaling—yet they often fail to reliably beat good old Chain of Thought (CoT) reasoning. So along with orchestration, the reasoning step itself warrants rethinking. 🔧 𝗢𝘂𝗿 𝗖𝗼𝗻𝘁𝗿𝗶𝗯𝘂𝘁𝗶𝗼𝗻: 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗥𝗲𝗮𝘀𝗼𝗻𝗶𝗻𝗴 𝗠𝗼𝗱𝘂𝗹𝗲 (𝗔𝗥𝗠) ARM is an agentic generalization of CoT, where each reasoning step is executed by an agentic workflow, automatically discovered via reflection-guided evolutionary search. 👉 Key ideas include: • 🧱 𝙎𝙘𝙖𝙛𝙛𝙤𝙡𝙙𝙚𝙙 𝙊𝙗𝙟𝙚𝙘𝙩𝙞𝙫𝙚: ARM is trained by replacing small windows of CoT traces, which solves 𝘤𝘳𝘦𝘥𝘪𝘵 𝘢𝘴𝘴𝘪𝘨𝘯𝘮𝘦𝘯𝘵 and improves 𝘱𝘦𝘳-𝘴𝘵𝘦𝘱 𝘤𝘰𝘮𝘱𝘦𝘵𝘦𝘯𝘤𝘦. • 🔀 𝘿𝙚𝙘𝙤𝙪𝙥𝙡𝙚𝙙 𝙏𝙧𝙖𝙞𝙣𝙞𝙣𝙜: A high-level Meta-Policy (MP) is trained separately to orchestrate step generator calls, showing reliable zero-shot transfer from the original CoT policy to ARM. • 🧾𝙀𝙢𝙥𝙞𝙧𝙞𝙘𝙖𝙡 𝙖𝙣𝙙 𝙏𝙝𝙚𝙤𝙧𝙚𝙩𝙞𝙘𝙖𝙡 𝙖𝙣𝙖𝙡𝙮𝙨𝙞𝙨: Under an MDP view, ARM’s stronger 𝘱𝘦𝘳-𝘴𝘵𝘦𝘱 𝘤𝘰𝘮𝘱𝘦𝘵𝘦𝘯𝘤𝘦 induces a policy improvement under mild CoT compatibility assumptions, giving bounded zero-shot transfer—verified through ablations. 💡 𝗞𝗲𝘆 𝗙𝗶𝗻𝗱𝗶𝗻𝗴𝘀: 1️⃣ 𝙎𝙮𝙨𝙩𝙚𝙢𝟣→𝙎𝙮𝙨𝙩𝙚𝙢𝟤: Improving the deductive reasoning step itself is more impactful than designing increasingly complex MAS orchestrations. 2️⃣ 𝙎𝙊𝙏𝘼 𝙋𝙚𝙧𝙛𝙤𝙧𝙢𝙖𝙣𝙘𝙚 𝘼𝙘𝙧𝙤𝙨𝙨 𝘿𝙤𝙢𝙖𝙞𝙣𝙨: Our combined ARM + MP system surpasses strong manual and automated baselines across math, science and reasoning benchmarks (𝘔𝘈𝘛𝘏𝟧𝟢𝟢, 𝘈𝘐𝘔𝘌, 𝘏𝘔𝘔𝘛, 𝘎𝘗𝘘𝘈, 𝘓𝘪𝘷𝘦𝘉𝘦𝘯𝘤𝘩) 3️⃣ 𝘾𝙧𝙤𝙨𝙨-𝙈𝙤𝙙𝙚𝙡 𝙍𝙤𝙗𝙪𝙨𝙩𝙣𝙚𝙨𝙨:ARM is a drop-in module for CoT, no domain or model specific tuning needed in our experiments. 🔍 𝗕𝗶𝗴 𝗣𝗶𝗰𝘁𝘂𝗿𝗲: This work suggests a shift in how we scale reasoning in LLMs: A system2 approach where each step is deliberately chosen opens a new frontier in inference time scaling. 📄 Full paper: https://lnkd.in/gNsqRaQN Grateful to my co-authors Bohan Yao and Vikas Yadav, for an amazing collaboration over the summer, and to ServiceNow CoreLLM team for their support and mentorship. Sagar Davasam, Sathwik Tejaswi Madhusudan, Srinivas Sunkara #NeurIPS #MATHAI #LLM #AIReasoning #MultiAgentSystems #AgenticAI
-
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
The paper addresses a foundational challenge in scientific modeling: how to build reliable mechanistic models of dynamical systems using large language models (LLMs). Mechanistic models are critical for understanding and simulating real-world processes in science and policy, but current LLM-based approaches often assume oversimplified conditions that don’t match practical settings—such as fully observed data or narrow task goals. To bridge this gap, the authors propose a new evaluation framework called Neural-Integrated Mechanistic Modeling (NIMM) that tests LLM-generated models under realistic conditions, including partial observations and a variety of task objectives. Their evaluation shows that existing baseline methods struggle across key dimensions like model effectiveness and correctness of generated code. In response, the paper introduces NIMMGen, an agentic framework that iteratively refines mechanistic models to improve both practical validity and code quality, addressing observed deficiencies in baselines. The authors demonstrate the utility of NIMMGen through experiments on three diverse scientific datasets, showing that it produces more accurate and usable mechanistic models than prior methods. Importantly, the learned models support counterfactual intervention simulation, meaning they can be used to explore “what-if” scenarios—a key capability for scientific insight and decision-making. By combining iterative refinement with neural integration, NIMMGen improves the reliability of mechanistic digital twins built with LLM assistance. These results suggest promising directions for building trustworthy, automated scientific models that maintain both mechanistic interpretability and data-driven flexibility. https://lnkd.in/gGcCg9_9
-
Furu Wei
Microsoft Research Asia • 13K followers
Introducing Generative Adversarial Distillation (GAD): a novel GAN-style formulation and framework that facilitates both on-policy and black-box distillation of large language models (LLMs). GAD is the first technique to enable block-box on-policy distillation from proprietary teachers where internal logits or parameters are inaccessible, or distillation between teacher and student LLMs with incompatible vocabularies. GAD expands our prior work on white-box on-policy distillation (i.e., MiniLLM), pioneering block-box on-policy distillation for LLM training. Specifically, GAD frames the student LLM as a generator and trains a discriminator to distinguish its responses from the teacher LLM’s, creating a minimax game. The discriminator acts as an on-policy reward model that co-evolves with the student, providing stable, adaptive feedback. Experimental results show that GAD consistently surpasses the commonly used sequence-level knowledge distillation. In particular, Qwen2.5-14B-Instruct (student) trained with GAD becomes comparable to its teacher, GPT-5-Chat, on the LMSYS-Chat automatic evaluation. The results establish GAD as a promising and effective paradigm for black-box LLM distillation. Our team has been conducting fundamental research in knowledge distillation with wide adoptions across the industry. - MiniLM: We introduced multi-head attention distillation, establishing the most effective distillation method for BERT-style models. The open-source MiniLM models (e.g., 6x384) have become the most widely utilized small encoder models on the Hugging Face. - MiniLLM: Our proposed Reverse KLD is recognized as one of the most effective, de facto on-policy distillation approaches for modern LLM training, which has been widely used by Thinking Machines, Gemma, and many other teams and models. - BitDistill: We proposed BitNet Distillation to finetune off-the-shelf full-precision LLMs (e.g., Qwen) into 1.58-bit precision (ternary weights {-1, 0, 1}), achieving performance parity with the full-precision counterparts on specific downstream tasks. - GAD: The development of Generative Adversarial Distillation (GAD) now allows for black-box on-policy distillation, overcoming two major prior limitations: (1) Distillation from proprietary teachers where internal logits or parameters are inaccessible; (2) Distillation between teacher and student LLMs with incompatible vocabularies. https://lnkd.in/gMaP2c7w
2 Comments