Note Probabl’s get together, in falls 2025 I’m thrilled to announce that I’m stepping up as Probabl ’s CSO (Chief Science Officer) to supercharge scikit-learn and its ecosystem, pursuing my dreams of tools that help go from data to impact. Scikit-learn, a central tool Scikit-learn is central …
At Maïc’s 100th birthday, I asked her “you lived 100 years, what was the most important change for you?”. She mentioned “Internet”. I asked, why was the Internet important to her eyes? Because this is how she kept close contact with her loved ones, sharing travels or discussing everyday …
I have recently been awarded France’s national order of merit , for my career, in science, in open source, and around AI. The speech that I gave carries messages important to me (French below; it flows better). Contents Speech translated to English Le texte d’origine, en Français
Note TabICL is a state-of-the-art tabular learner [Qu et al 2025] . The key is its very rich prior, that is baked in a pre-trained architecture -a table foundation model-, and leveraged by in-context-learning. Thanks to clever choices, it is fast and scalable, efficient even without a GPU. Contents Recent progress …
Note For me, 2024 was full of back and forth between research, software, and connecting these to society. Here, I lay out some highlights on AI and society, as well as research and software, around tabular AI and language models. As 2025 starts, I’m looking back on 2024. It …
Note Foundation models, pretrained and readily usable for many downstream tasks, have changed the way we process text, images, and sound. Can we achieve similar breakthroughs for tables? Here I explain why with “CARTE” , we’ve made significant headway. Contents Pre-training for data tables: hopes and challenges Pre-training is a …
We just released skrub 0.2.0 . This release markedly simplifies learning on complex dataframes. model = tabular_learner(‘classifier’) Simple, yet solid default baseline The highlight of the release is the tabular_learner function, which facilitates creating pipelines that readily perform machine learning on dataframes, adding preprocessing to a scikit-learn compatible learner …
Note Open-source efforts around scikit-learn at Inria are spinning off to a new enterprise, Probabl , in charge of sustainable development of a data-science commons. Contents Prelude: funding scikit-learn is hard The birth of a new ambition Probabl, a mission-driven enterprise Probabl is already having an impact My position within Probabl …
Note François Chollet rightfully said that people often underestimate the impact of scikit-learn. I give here a few illustrations to back his claim. A few days ago, François Chollet (the creator of Keras, the library that that democratized deep learning) posted : Indeed, scikit-learn continues to be the most popular machine …
English summary I have been appointed to the government-level panel of experts on AI, to set the national vision and strategy in France. J’ai l’honneur d’être nommé au comité de l’intelligence artificielle du gouvernement Français . La mission qui nous est confiée d’éclairer l’action publique …
A retrospective on last year (2022): I embarked on a new scientific adventure, assembling a team focused on developing machine learning for health and social science. The team has existed for almost a year, and the vision is nice shaping up. Let me share with you illustrations of where we …
The Mayavi Python software, and my personal history: A thread on Python and scipy ecosystems, building open source codebase, and meeting really cool and friendly people I am writing today as a goodbye to the project: I used to be one of the core contributors and maintainers but have been …
Broad decoding models that can specialize to discriminate closely-related mental process with limited data TL;DR Decoding models can help isolating which mental processes are implied by the activation of given brain structures. But to support a broad conclusion, they must be trained on many studies, a difficult problem given …
Note Join us to work on reinventing data-science practices and tools to produce robust analysis with less data curation. It is well known that data cleaning and preparation are a heavy burden to the data scientist. Dirty data research In the dirty data project , we have been conducting machine-learning research …
Note With the growth of scikit-learn and the wider PyData ecosystem, we want to recruit in the Inria scikit-learn team for a new role . Departing from our usual focus on excellence in algorithms, statistics, or code, we want to add to the team someone with some technical understanding, but an …
The year 2020 has undoubtedly been interesting: the covid19 pandemic stroke while I was on a work sabbatical in Montréal, at the MNI and the MILA , and it pushed further my interest in machine learning for health-care. My highlights this year revolve around basic and applied data-science for health . Highlights …
Note This post discuss the difficulties of communicating while developing open-source projects and tries to gives some simple advice. A large software project is above all a social exercise in which technical experts try to reach good decisions together, for instance on github pull requests. But communication is difficult, in …
Jean Dechoux was born between the first and the second world wars, in a small French town, close to Germany. His family was that of poor farmers, who would work in coal mines to make up for the small size of their crops. He grew to become a pulmonologist, heading …
Note A simple survey asking authors of two leading machine-learning conferences a few quantitative questions on their experimental procedures. How do machine-learning researchers run their empirical validation? In the context of a push for improved reproducibility and benchmarking, this question is important to develop new tools for model comparison. We …
My current research spans wide: from brain sciences to core data science. My overall interest is to build methodology drawing insights from data for questions that have often been addressed qualitatively. If I can highlight a few publications from 2019 [1] , the common thread would be computational statistics, from dirty …
Note Given two set of observations, are they drawn from the same distribution? Our paper Comparing distributions: l1 geometry improves kernel two-sample testing at the NeurIPS 2019 conference revisits this classic statistical problem known as “two-sample testing”. This post explains the context and the paper with a bit of hand …
Note An important acknowledgement for a different view of doing science: open, collaborative, and more than a proof of concept. A few days ago, Loïc Estève, Alexandre Gramfort, Olivier Grisel, Bertrand Thirion, and myself received the “Académie des Sciences Inria prize for transfer” , for our contributions to the scikit-learn project …
From a scientific perspective, 2018 [1] was once again extremely exciting thank to awesome collaborators (at Inria , with DirtyData , and our local scikit-learn team ). Rather than going over everything that we did in 2018, I would like to give a few highlights: We published major work using machine learning to …
We have just announced that a foundation will be supporting scikit-learn at Inria [1] : scikit-learn.fondation-inria.fr Growth and sustainability This is an exciting turn for us, because it enables us to receive private funding. As a result, we will be able to have secure employment for some existing core …
Two weeks ago, we held a scikit-learn sprint in Austin and Paris. Here is a brief report, on progresses and challenges. Several sprints We actually held two sprint in Austin: one open sprint, at the scipy conference sprints, which was open to new contributors, and one core sprint, for more …
In my opinion the scientific highlights of 2017 for my team were on multivariate predictive analysis for brain imaging: a brain decoder more efficient and faster than alternatives, improvement clinical predictions by predicting jointly multiple traits of subjects, decoding based on the raw time-series of brain activity, and a personnal …
Note Scientific progress calls for reproducing results. Due to limited resources, this is difficult even in computational sciences. Yet, reproducibility is only a means to an end. It is not enough by itself to enable new scientific results. Rather, new discoveries must build on reuse and modification of the state …
Two week ago, we held in Paris a large international sprint on scikit-learn . It was incredibly productive and fun, as always. We are still busy merging in the work, but I think that know is a good time to try to summarize the sprint. A massive workforce We had a …
Year 2016 has been productive for science in my team . Here are some personal highlights: bridging artificial intelligence tools to human cognition, markers of neuropsychiatric conditions from brain activity at rest, algorithmic speedups for matrix factorization on huge datasets… Artificial-intelligence convolutional networks map well the human visual system Eickenberg et …
To my friends developing data science for the social media, marketing, and advertising industries, It is time to accept that we have our share of responsibility in the outcome of the US elections and the vote on Brexit. We are not creating the society that we would like. Facebook, Twitter …
At MLOSS 15 we brainstormed on reproducible science, discussing why we care about software in computer science . Here is a summary blending notes from the discussions with my opinion. “Without engineering, science is not more than philosophy” — the community How do we enable better Science? Why do we do software …
Nilearn’s goals Make advanced machine learning techniques easy for neuroimaging research. After 6 months of efforts, We just released version 0.2 of nilearn , dedicated to making machine learning in neuroimaging easier and more powerful . This release integrates the features of the july sprint , and more . Highlights Better documentation …
My research group is looking to fill a post-doc position on learning biomarkers from functional connectivity . Scientific context The challenge is to use resting-state fMRI at the level of a population to understand how intrinsic functional connectivity captures pathologies and other cognitive phenotypes. Rest fMRI is a promising tool for …
Note The 2015 edition of the machine learning open source software (MLOSS) workshop was full of very mature discussions that I strive to report here. I give links to the videos. Some machine-learning researchers have great thoughts about growing communities of coders, about code as a process and a deliverable …
A couple of weeks ago, we had in Paris the second international nilearn sprint, dedicated to making machine learning in neuroimaging easier and more powerful . It was such a fantastic experience, as nilearn is really shaping up as a simple yet powerful tool, and there is a lot of enthusiasm …
Note tl;dr: Reproducibilty is a noble cause and scientific software a promising vessel. But excess of reproducibility can be at odds with the housekeeping required for good software engineering. Code that “just works” should not be taken for granted. This post advocates for a progressive consolidation effort of scientific …
Note This year again we will have an exciting workshop on the leading-edge machine-learning open-source software. This subject is central to many, because software is how we propagate, reuse, and apply progress in machine learning. Want to present a project? The deadline for the call for papers is Apr 28th …
We, Parietal team at INRIA , are recruiting software developers to work on open source machine learning and neuroimaging software in Python. In general, we are looking for people who: have a mathematical mindset, are curious about data (ie like looking at data and understanding it) have an affinity for problem-solving …
EuroScipy 2015, the annual conference on Python in science will take place in Cambridge, UK on 26-30 August 2015. The conference features two days of tutorials followed by two days of scientific talks & posters and an extra day dedicated to developer sprints. It is the major event in Europe in …
I am moving my website to a new design, relying on Pelican and more modern CSS. So far, I had been using rest2web to generate the static part of the website, and a local install of wordpress for the blog. I wasn’t doing good on keeping the wordpress install …
Work with us to leverage leading-edge machine learning for neuroimaging At Parietal , my research team, we work on improving the way brain images are analyzed, for medical diagnostic purposes, or to understand the brain better. We develop new machine-learning tools and investigate new methodologies for for quantifying brain function from …
A week ago, the 2014 edition of the scikit-learn sprint was held in Paris. This was the third time that we held an internation sprint and it was hugely productive, and great fun, as always. Great people and great venues We had a mix of core contributors and newcomers, which …
We have just released the 0.15 version of scikit-learn. Hurray!! Thanks to all involved . A long development stretch It’s been a while since the last release of scikit-learn . So a lot has happened. Exactly 2611 commits according my count. Quite clearly, we have more and more existing code …
I’d like to welcome the four students that were accepted for the GSoC this year: Issam: Extending Neural networks Hamzeh: Sparse Support for Ensemble Methods Manoj: Making Linear models faster Maheshakya: Locality Sensitive Hashing Welcome to all of you. Your submissions were excellent, and you demonstrated a good will …
Work with us on putting machine learning in the hand of cognitive scientists Parietal is a research team that creates advanced data analysis to mine functional brain images and solve medical and cognitive science problems. Our day to day work is to write machine-learning and statistics code to understand and …
Christophe Pradal, Hans Peter Langtangen, and myself recently edited a version of the Journal of Computational Science on scientific software, in particular those written in Python. We wrote an editorial defending writing and publishing open source scientific software that I wish to summarize here. The full text preprint is openly …