RSSAmplifier

Blog

Konrad Hinsen's blog

blog.khinsen.netRSS feed ↗94 posts

Latest posts

Conviviality in computational science

Convivial technology was defined by Ivan Illich in his 1973 book "Tools for conviviality" as technology that supports a convivial society, which is a society that strives to grant each of its members as much agency as is possible without infringing on other members agency. Conviviality is thus about equality, about the absence of dominance relations. Convivial technology is shaped by its users…

Cultures of making and relating

Cultures of Programming - The Development of Programming Concepts and Methodologies is a recent book by Tomáš Petříček that analyses the history of programming from the perspective of five interwoven cultures. It contains a lot of interesting insight, so I encourage you to read it. At the very least, read the first chapter. In this post, I try to relate these five cultures to the wider world of…

Automating science

The advent of AI agents based on large language models (LLMs) has put the idea of automating the intellectual and cognitive work of researchers on the table. A lively, sometimes even heated discussion is already going on. A frequently missing piece in this debate is the question why we, individually and as a society, actually do science. I will examine this question first, and then consider what…

Preparing for scientific deepfakes

By now, most scientists have probably seen figures, tables, and even entire journal articles made by so-called "generative AI", containing more or less subtle mistakes or inconsistencies. What I haven t seen yet, but expect to see soon, is the scientific equivalent of deepfakes: made-up results that come with made-up code that reproduces them. This is likely to become a new challenge for…

Explorable explorable explanations

A much cited essay by Bret Victor, "Explorable Explanations" , argues for supporting and encouraging active reading in communicating ideas. Explanatory text should thus be complemented by interactive visualizations and computational demonstrations, allowing the reader to actively engage with the ideas. If you haven t read Victor s essay yet, please do so now, and then come back here. It s not very…

Explaining software and computational methods

How can we document software and computational analyses in such a way that others can convince themselves of their validity, and build on them for their own work? The question has been around for many years, and a number of attempts have been made to provide partial answers. This post provides a brief review and describes my own tentative answer, inviting you to play with it. Explainable AI is a…

Why computational reproducibility matters

Thirty years after my first contact with computational (ir)reproducibility, I am happy to note that many things have improved. Reproducibility, computational and otherwise, is increasingly recognized as an important aspect of scientific quality control, and mostly considered worth striving for. However, I also note that more and more people, including reproducibility activists, have lost contact…

Why we should review research software

At the recent SciCodes Symposium , I brought up the question of reviewing research software during the panel discussion. One panelist then raised the question of why we should review research software. I found this question surprising at first, but I do agree that it deserves an answer. Here is mine. My goal with bringing up the question was to learn out the state of the art: which institutions,…

Going for robustness: science

This is a follow-up to my earlier post entitled "Going for robustness" , focusing on scientific research. What is "robust science"? I see at least two interpretations, and I am going to discuss both of them: robustness of scientific findings, and robustness of the process of doing science, which includes in particular the robustness of the web of scientific research institutions: first and…

Going for robustness

I suspect that most people in the Western world (at least) are realizing that we are living in interesting times . News of floodings, droughts, and wildfires are ever more frequent. We hear that this is due to climate change, which most governments promise to fight, but don t. Our economies keep growing, but our quality of life is not improving. Digital tools are ever more prominent in our lives,…

Modular malleability

This is a contribution to the Challenge problem: Fearless extensibility by the Malleable Systems Collective . As superheroes know very well, with great power comes great responsibility . Malleable systems offer a lot of power to their users. In exchange, users have to take responsibility for the system they have tailored to their needs, since there is nobody else to blame. If you are the only user…

The low-hanging fruit in computational reproducibility

Yesterday I participated in the International workshop “Software, Pillar of Open Science” , organized by the French Committee for Open Science . In the course of the various presentations and discussions (both in public and during coffee breaks), I realized that something has been absent from such events all the time: the vast majority of scientists. What prompted this insight was the…

This blog gets a facelift

Regular visitors to my blog have probably noticed that it looks different now. However, the visual changes are only a side effect of a more profound change: I now use a different static site generator, coleslaw . It s been a while that I wanted to replace Disqus by a less invasive commenting system, and the recent announcement by Disqus to insert ads into the comments on my blog was what finally…

Following branching conversations on Mastodon

This post is a follow-up to my previous one, Deconstructing the Mastodon client . My topic is a scenario that traditional Mastodon clients handle rather badly, wheres my home-grown solution handled it very well: lengthy and branching conversations. Such conversations happen all the time on social networks. Someone posts an interesting question or observation, which is commented by many others.…

Deconstructing the Mastodon client

Ever since I joined Twitter in 2011, and then moved to Mastodon in 2022, I have been unhappy with the timeline view proposed by both of these communication platforms as their main interface. Now I have finally done something about it: I wrote my own Mastodon client. Or perhaps rather a non-client, because the concept of "the client" is a big part of what I disliked. My use of social networks can…

Welcome to my digital garden!

A few years ago, I discovered Mike Caulfield s The Garden and the Stream: A Technopastoral and understood why I wasn t happy with my blog. Blogs are streams, timelines of posts. Each post has a timestamp, and is considered "finished". Later changes are technically possible, but culturally limited to corrections. A blog post is considered a published essay, and therefore comes with a date of…

The dependency hubs in Open Source software

A few days ago, Google announced its experimental project Open Source Insights , which permits the exploration of the dependency graph of Open Source software. My first look at it ended with a disappointment: in its initial stage, the site considers only the package universes of Java, JavaScript, Go, and Rust. That excludes most of the software I know and use, which tends to be written mainly in…

The structure and interpretation of scientific models, part 2

In my last post , I have discussed the two main types of scientific models: empirical models, also called descriptive models, and explanatory models. I have also emphasized the crucial role of equations and specifications in the formulation of explanatory models. But my description of scientific models in that post left aside a very important aspect: on a more fundamental level, all models are…

The structure and interpretation of scientific models

It is often said that science rests on two pillars, experiment and theory. Which has lead some to propose one or two additional pillars for the computing age: simulation and data analysis. However, the real two pillars of science are observations and models. Observations are the input to science, in the form of numerous but incomplete and imperfect views on reality. Models are the inner state of…

Some comments on AlphaFold

Many people are asking for my opinion on the recent impressive success of AlphaFold at CASP14 , perhaps incorrectly assuming that I am an expert on protein folding. I have actually never done any research in that field, but it s close enough to my research interests that I have closely followed the progress that has been made over the years. Rather than reply to everyone individually, here is a…

The four possibilities of reproducible scientific computations

Computational reproducibility has become a topic of much debate in recent years. Often that debate is fueled by misunderstandings between scientists from different disciplines, each having different needs and priorities. Moreover, the debate is often framed in terms of specific tools and techniques, in spite of the fact that tools and techniques in computing are often short-lived. In the…

The landscapes of digital scientific knowledge

Over the last years, an interesting metaphor for information and knowledge curation is beginning to take root. It compares knowledge to a landscape in which it identifies in particular two key elements: streams and gardens. The first use of this metaphor that I am aware of is this essay by Mike Caulfield , which I strongly recommend you to read first. In the following, I will apply this metaphor…

An open letter to software engineers criticizing Neil Ferguson's epidemics simulation code

Dear software engineers, Many of you were horrified at the sight of the C++ code that Neil Ferguson and his team wrote to simulate the spread of epidemics . I feel with you. The only reason why I am less horrified than you is that I have seen a lot of similar-looking code before. It is in fact quite common in scientific computing, in particular in research projects that have been running for many…

Wanted: a hierarchically modular software architecture

In his 1962 classic "The Architecture of Complexity" , Herbert Simon described the hierarchical structure found in many complex systems, both natural and human-made. But even though complexity is recognized as a major issue in software development today, the architecture described by Simon is not common in software, and in fact seems unsupported by today s software development and deployment…

Emacs as a malleable system

Malleable systems are software systems that are designed to be modified and extended by their users, eliminating the usually strict borderline between developers and users. Making scientific software more malleable is a goal that I have been pursuing for 25 years, starting with a shift from Fortran to Python as my main programming language, and a simultaneous shift from writing programs to writing…

The rise of community-owned monopolies

One question I have been thinking about in the context of reproducible research is this: Why is all stable software technology old, and all recent technology fragile? Why is it easier to run 40-year-old Fortran code than ten-year-old Python code? A hypothesis that comes to mind immediately is growing code complexity, but I d expect this to be an amplifier rather than a cause. In this pose, I will…

Pharo year one

It s the season when everyone writes about the past year, or even the past decade for a year number ending in 9. I ll make a modest contribution by summarizing my experience with Pharo after one year of using it for projects of my own. My first contact with Pharo happened a bit more than one year ago, when I signed up for the Pharo MOOC in October 2018. But following a MOOC means working on…

Industrialization of scientific software: a case study

A coffee break conversion at a scientific conference last week provided an excellent illustration for the industrialization of scientific research that I wrote about in a recent blog post . It has provoked some discussion on Twitter that deserves being recorded and commented on a more permanent medium. Which is here. I was chatting with a colleague who I have been meeting at such occasions for…

The industrialization of scientific research

Over the last few years, I have spent a lot of time thinking, speaking, and discussing about the reproducibility crisis in scientific research. An obvious but hard to answer question is: Why has reproducibility become such a major problem, in so many disciplines? And why now? In this post, I will make an attempt at formulating an hypothesis: the underlying cause for the reproducibility crisis is…

The computational notebook of the future (part 2)

A while ago I wrote about my ideas for a successor of today s computational notebooks. Since then I have made some progress on a prototype implementation, which is the topic of this post. Again I have made a companion screencast so that you can get a better idea of how all this works in practice. As a reminder, the two aspects of today s notebooks ( Mathematica , Jupyter , R markdown ,…

Is reproducibility good for scientific progress? (a paper review)

A few days ago, a discussion in my Twitter timeline caught my attention. It was about a very high-level model for the process of scientific research whose conclusions included the affirmation that reproducibility does not improve the convergence of the research process towards truth. The Twitter discussion set off some alarm bells for me, in particular the use of the term "reproducibility" in the…

The computational notebook of the future

Regular readers of this blog may have noticed that I am not very happy with today s state of computational notebooks, such as they were pioneered by Mathematica and popularized by more recent free incarnations such as Jupyter , R markdown , or Emacs/OrgMode . In this post and the accompanying screencast (my first one!), I will explain what I dislike about today s notebooks, and how I think we can…

Exploring Pharo

One of the more interesting things I have been playing with recently is Pharo , a modern descendent of Smalltalk. This is a summary of my first impressions after using it on a small (and unfinished) project , for which it might actually turn out to be very helpful. The first time I read about Smalltalk was in the August 1981 issue of Byte magazine . Back then, I was a high school student and I had…

Knowledge distillation in computer-aided research

There is an important and ubiquitous process in scientific research that scientists never seem to talk about. There isn t even a word for it, as far as I now, so I ll introduce my own: I ll call it knowledge distillation . In today s scientific practice, there are two main variants of this process, one for individual research studies and one for managing the collective knowledge of a discipline. I…

Literate computational science

Since the dawn of computer programming, software developers have been aware of the rapidly growing complexity of code as its size increases. Keeping in mind all the details in a few hundred lines of code is not trivial, and understanding someone else s code is even more difficult because many higher-level decisions about algorithms and data structures are not visible unless the authors have…

Scientific software is different from lab equipment

My most recent paper submission ( preprint available) is about improving the verifiability of computer-aided research, and contains many references to the related subject of reproducibility. A reviewer asked the same question about all these references: isn t this the same as for experiments done with lab equipment? Is software worse? I think the answers are of general interest, so here they are.…

Scientific communication is a research problem

A recent article in "The Atlantic" has been the subject of many comments in my Twittersphere. It s about scientific communication in the age of computer-aided research, which requires communicating computations (i.e. code, data, and results) in addition to the traditional narrative of a paper. The article focuses on computational notebooks, a technology introduced in the late 1980s by Mathematica…

What can we do to check scientific computation more effectively?

It is widely recognized by now that software is an important ingredient to modern scientific research. If we want to check that published results are valid, and if we want to build on our colleagues published work, we must have access to the software and data that were used in the computations. The latest high-impact statement along these lines is a Nature editorial that argues that with any…

Data science in ancient Greece

Data science is usually considered a very recent invention, made possible by electronic computing and communication technologies. Some consider it the fourth paradigm of science, suggesting that it came after three other paradigms, though the whole idea of distinct paradigms remains controversial. What I want to point out in this post is that the principles of data science are much older than most…

Stability in the SciPy ecosystem: a summary of the discussion

The plea for stability in the SciPy ecosystem that I posted last week on this blog has generated a lot of feedback, both as comments and in a lengthy Twitter thread . For the benefit of people discovering it late, here is a summary of the main arguments and my reply to them. Just freeze your code and it will be reproducible forever By far the most frequent argument against my claim that we need…

A plea for stability in the SciPy ecosystem

Two NumPy-related news items appeared on my Twitter feed yesterday, just a few days after I had accidentally started a somewhat heated debate myself concerning the poor reproducibility of Python-based computer-aided research. The first was the announcement of a plan for dropping support for Python 2 . The second was a pointer to a recent presentation by Nathaniel Smith entitled "Inside NumPy" and…

There is no such thing as software development

It s hard to find an aspect of modern life that is not influenced in some way by software. Some of it is very visible, for example the Web browser I start on my computer. Other software is completely invisible, such as the software controlling my car s diesel engine. Some software is safety critical, for example flight control software in airplanes. Other software is used in a much more futile…

Why Python does so well in scientific computing

A few days ago, I noticed this tweet in my timeline: I 'still' program in C. Why? Hint: it's not about performance. I wrote an essay to elaborate... appearing at Onward! https://t.co/pzxjfvUs5B — Stephen Kell (@stephenrkell) September 5, 2017 That sounded like a good read for the weekend, which it was. The main argument the author makes is that C remains unsurpassed as a system integration…

Which mistakes do we actually make in scientific code?

Over the last few years, I have repeated a little experiment: Have two scientists, or two teams of scientists, write code for the same task, described in plain English as it would appear in a paper, and then compare the results produced by the two programs. Each person/team was asked to do a maximum amount of verification and testing before comparing to the other person s/team s work. Let me state…

Reproducible research in the Python ecosystem: a reality check

A few years ago, I decided to adopt the practices of reproducible research as far as possible within the technical and social constraints I have to live with. So how reproducible is my published code over time? The example I have chosen for this reproducibility study is a 2013 paper about computing diffusion coefficients from molecular simulations. All code and data has been published as an…

Reproducibility does not imply reproduction

In discussions about computational reproducibility (or replicability, or repeatability, according to the preference of each author), I often see the argument that reproducing computations may not be worth the investment in terms of human effort and computational resources. I think this argument misses the point of computational reproducibility. Obviously, there is no point in repeating a…

Sustainable software and reproducible research: dealing with software collapse

Two currently much discussed issues in scientific computing are the sustainability of research software and the reproducibility of computer-aided research. I believe that the communities behind these two ideals should work together on taming their common enemy: software collapse. As a starting point, I propose an analysis of how the risk of collapse affects sustainability and reproducibility. What…

From reproducible to verifiable computer-aided research

The importance of reproducibility in computer-aided research (and elsewhere) is by now widely recognized in the scientific community. Of course, a lot of work remains to be done before reproducibility can be considered the default. Doing computational research reproducibly must become easier, which requires in particular better support in computational tools. Incentives for working and publishing…

Composition is the root of all evil

Think of all the things you hate about using computers in doing research. Software installation. Getting your colleagues scripts to work on your machine. System updates that break your computational code. The multitude of file formats and the eternal need for conversion. That great library that s unfortunately written in the wrong language for you. Dependency and provenance tracking.…

On HDF5 and the future of data management

Yesterday a blog post by Cyrille Rossant entitled "Moving away from HDF5" caught my eye. My own tendency at the moment is to use HDF5 more and more, so I was interested in why someone else would want to do the opposite. Here is my conclusion after reading his post, plus some ideas about where scientific data management is or should be heading in my opinion. Any evaluation of some technology…