RSS Amplifier

Dr Jo · Aug 16, 2026

Counting from Zero

0
Sign in to vote or save

Dr Jo · Dr Jo

“First Thoughts are the everyday thoughts. Everyone has those. Second Thoughts are the thoughts you think about the way you think. People who enjoy thinking have those. Third Thoughts are thoughts that watch the world and think all by themselves. They’re rare, and often troublesome. Listening to them is part of witchcraft.”
Terry Pratchett, A Hat Full of Sky.

Of course the title is a metaphor! That’s the whole point.

This series of posts is about the deeper aspects of programming. If you believe that all that is now needed to program is to ask an ‘AI’ or large language model to compose code—well then, Boy have you come to the wrong place.

No prior knowledge is needed. We’re going to start at the very beginning, and move slowly. Right from that beginning, we’re going to look at metacognition: thinking about thinking. Third Thoughts may even intrude, on occasion. But our enquiry will be intensely practical: over time, we will also build an entire, new, fully functional language from the ground up. Within a few posts, I’ll introduce her; there’s no need to rush here. Exploration often takes time.

I can guarantee two things. One is that you have never met this language before. The other is that she is substantially unlike any other programming language in existence. Neither may necessarily be a good thing. Or, for that matter, a bad thing. Just different. Let’s see.

About half a century ago—God, I’m old!—I started asking questions about computers. I haven’t stopped. It all began in my first year at medical school. I had a choice, you see. I could either take a half-course in sociology, or a half-course in computer programming.

The lady who ran the sociology course was a haughty German with a lisp,1 so I chose programming instead. Ironically, although the course involved inputting simple FORTRAN IV programs on punch cards, I soon turned to the Lisp programming language, so it seems I couldn’t escape the lisp.

My observation then—even with the FORTRAN—was “Gosh! We can teach these machines to think!” My immediate question was “How do we do this right?” But this speculation led in turn to other questions like “What is a good language?” and even “How do we build one?” These are the difficult questions I’m going to try to answer in subsequent posts. We will start very simply, of course. With counting.

To anyone who can count in a natural language like English or Chinese, Arabic or Nasioi, two things are obvious. The first is that counting is, well, simple and natural. The second is that counting in every other language is just unnatural and wrong.

Ethnomathematicians—people who study how different cultures count and do mathematics—have had a lot of fun contrasting those two observations. Obviously, you need to take some time to study those various languages, and time is running out for many of them. They are being supplanted at an increasing pace by just a few, dominant newcomers.

Unequivocally the greatest source of diverse languages is Papua New Guinea (PNG), where languages likely extend back for over 40,000 years. The late Glendon Lean studied counting systems there for two decades, documenting details of over 1,500 counting systems in 883 languages.

He found a huge variety, almost none of them founded on base ten. Many worked in cycles, for example you might have a combination of a 5-cycle and a 20-cycle, with separate words for 1–5, and 20. You shouldn’t find this too strange, given that every time you tell the time of day, you’re working with an Egyptian cycle of 12 (or 2×12) hours, and Babylonian cycles of 60 for minutes and seconds.2

Other counting systems encountered by Lean used body-part tally systems to form a cycle, the most common being 27. One body-part tally system is shown above. Different classes of objects may be counted using different systems. The Kewa combine a 47-cycle body-part tally system with a 4-cycle system. Multiplication and addition can be combined, for example the Ndom counting system has 2 as ‘thef’ and 6 as ‘mer’. Eight is mer abo thef (6+2); 12 is mer an thef (6×2). In Yu-Wooi, 27 is “angek yem yemsi simb yem yemsi, angek yemsi, tak”, literally both hands, both legs, half hand and two.

But my favourite is the counting system of the Nasioi from Bougainville. The basics are well described in Conrad Hurd’s paper Nasioi Projectives, from Oceanic Linguistics volume 16(2) 111–178. In order to count something, you need to modify the ‘count’ based on what you’re counting. You use different words for counting different classes of objects—for example dance troupes (or bundles of bamboo), grandchildren of a common grandmother, houses (or tens of opossums), or ropes.3 There’s a lot more:

This paper describes a genre of words in the Nasioi language whose common denominator is that they obligatorily take suffixes of case-gender-number, called CGNs in this paper. At a loss for a name for these words, I have merged the term pronoun and adjective into projective, since most of them can stand for nouns, modify nouns, and be modified by nouns.

… Altogether, the projective genre includes, among other things, about 100 counting systems, 100 definite articles, 100 interrogative pronouns, 900 possessive pronouns, and 1600 demonstrative nouns.

A sort of “counting web”. It’s almost like object-oriented programming in its complexity :)

At this point you might feel that this has nothing to do with computer programming.4 Simplistically, you may be right. However, in terms of the meta-cognition required to program well, the idea that different people see a ‘simple’ concept like ‘numbers’ differently has everything to do with programming!

We will discover—repeatedly—that our pre-conceptions about how things should work get in the way of getting things to work well. It is extraordinarily difficult to shake off the mental shackles of a very specific point of view. If anything, this is becoming more and more difficult, whether we’re dealing with ethnomathematics, programming languages, or the homogenisation of thought we get with AI slop.

I first came across Glendon Lean’s work when I downloaded large parts of his 1992 PhD thesis from a website in PNG about 20 years ago. That website is no longer functioning, and if I want the same today, I am compelled to go behind a corporate paywall. Heaven help someone in PNG trying to recover their heritage.5

One of Lean’s most salient observations on the mathematical systems used in hundreds of different languages, and their different ways of viewing reality, was that as he was documenting so they were dying—both the old, fluent speakers, and the languages themselves. They have been eradicated by the intrusion of a very small number of more homogeneous languages.

I’m sure several of my readers will now raise their eyebrows and say something like this: “So what! These are just primitive number systems, and our current ones are manifestly better. There’s no loss here.”

To which the only reasonable response is a question “How can you be so sure?” The obvious imputation is that either the speaker has carefully studied these languages, and weighed them up, and compared them among one another and then with modern counting systems—or they are simply assuming that their way is better, on pretty much no basis at all other than ‘success’, in other words ‘eradication’. Which would be a bit silly, wouldn’t it? We’ll repeatedly discover how daft this sort of assumption can be. It’s like assuming the ‘cave art’ from Altamira has nothing to teach us, because it was made tens of thousands of years ago. In fact, it is an invaluable resource, because we can often see how extinct species looked, based on observations by contemporary observers. It’s also f-cking awesome.

If you visit this obscure Brazilian link you may (if it’s still up) be able to view a paper by Kay Owens and Charly Muke: Revising the history of number: how Ethnomathematics transforms perspectives on indigenous cultures.

Again, you may find this irrelevant—until you work out that most choices we have made in modern programming are cultural artefacts, rather than being ‘optimal design’. (This observation will make a few programmers so angry that they will refuse to continue, which is perhaps just as well.)

If we go back to the middle of the nineteenth century, the European view of the world was that counting systems and numerals that supported them were hallmarks of ‘civilisation’, and that these originated about 4000 BCE in either Egypt or Sumer. Counting then diffused elsewhere, throughout the world, from these cradles of civilisation—culminating in the acme of civilisation, which was, of course, to their minds European. Even ‘primitive’ societies that could count based on a “pure two cycle system”6 likely owed this ability to early diffusion of counting concepts from these cradles. This has been called “diffusionist theory”. It also posited that various perversions of a natural ‘1,2’ and base ten system accounted for all the ‘anomalous’ systems encountered around the world.

This theory is, of course, almost complete bullshit. We are pretty certain that languages have been evolving in different directions in PNG for at least 40,000 years. Comparative linguistic methods struggle to go back much beyond about 6–10,000 years, but it’s likely that many PNG languages and their counting systems diverged way before then—often with a lot of subsequent cross-pollination. We can still study the variation. Lean found almost no “pure two-cycle systems” in PNG.7 He did document the enormous variety we’ve hinted at, including body tally systems and cycles. There are even six-cycles (e.g. Kanum from Kolopom Island).

There’s a hilarious counterpoint here. Ask a programmer “How many bytes in a kilobyte?” The answer is, of course, 1024, rather than 1000, because 210 is 1024. In Indonesia near the border of West Papua and PNG, the number ‘ntamnao’ is 1296. You can work out this is 64—but according to Owens & Muke, younger Indonesians are now interpreting ntamnao as ‘1000’, referring to the currency note of that denomination. I guess this is a cultural parallel to “a bit more than 1000 bytes”.

There are 2 truly difficult problems in Computer Science:

0: Naming things

1: Cache invalidation

2: Off by one errors” — Reddit (Jokes)

In programming, certain themes come up again and again. Like “just getting it wrong by 1.”8 There are also different counting systems, and cultural tendencies that are no less strange than counting in Nasioi. One prevalent theme that you encounter among novice programmers is irritation that “counting begins at zero”.

If you program in Python, C, or C++, Java, JavaScript, Rust, Go, Perl, PHP, Ruby, Dart, Kotlin, C# and many other languages, you count from zero.

Starting at zero is pretty counter-intuitive, isn’t it? You’ll see that the sub-title of my post is “How to Program, part 1/lots” and not “How to Program, part 0/lots”. This is in deference to common usage, so why ever start at zero?

Unless, of course, the programming language you’re using was largely invented by mathematicians, in which case counting starts at 1. This can be maddening to other programmers. Languages like R, Fortran, Julia, MATLAB, and dinosaurs like PL/I and APL.9

We can understand the mathematical aversion to zero, but why should programmers even consider starting there? The answer is in the indexing. If you have some way of storing a sequence of numbers in computer memory, along these lines:

3 | 1 | 4 | 1 | 5 | 9 | 2 | 6 | 5 | 3 | 5 …

… then it would be neat to have some way of saying “Give me the nth number”. But if we conceptually “Ask a question” like:

… then what is AT “Position 4” in that ‘list’ or ‘array’? Do we return a “1” or a “5”? That depends on your language, doesn’t it?

To understand better, we need to dig in a bit more. First, I need to confess. I snuck in a little implicit assumption there. Did you spot it? Yep, I talked about “storing a sequence of numbers” without even hinting how those numbers might be represented. For convenience, I plonked them in parenthesis, with spaces as separators, and called them a ‘list’ or an ‘array’. But this invites all sorts of questions. Questions like:

  • “How big a number can I represent/store at a given position?”

  • “What sort of number are we talking about? Are we limited to integers?”

  • “Can I store negative numbers, and if so, how are these represented internally?”

  • “You said ‘store’. How do we actually move numbers around?”

  • “How persistent are those numbers? Will some future spelunker be able to read them on a cave wall in 40,000 years’ time?”

Can you now see how important apparently risible ideas like ‘cultural context’ are for numbers? In future posts, we’ll try to answer all of those questions, but for now let’s assume that you can store integers in memory, up to a certain size, including the number zero.

We will also assume there is a way we can pick up that number (Let’s call this temporary billet for a number we’ve just picked up a ‘register’) and later write the value in the register back to a position (index) in memory. But how do we specify that position? It seems logical to use another register, and even more logically, we can call this an “index register”.

So why might it be logical to start counting at zero? Yep. If we have a zero in that index register, it makes enormous sense to regard that as the initial storage position in memory. Index zero. Otherwise we have to “artificially increment by one”, simply to satisfy our mathematical, intuitive sense as to where counting starts. Zero is wasted. Cue the angry mathematician, who simply has to design her COBOL (or his Fᴏʀᴛʀᴀɴ) system differently.

It just makes sense—but it just makes sense depending on your cultural bent.

We have just got started. The cultural hilarity that attaches to counting in English, and ultimately doing other arithmetic in English—all things that we might see as entirely natural—will need to wait for my next post. In fact, it’s all so crazy that we may need several posts to get to a point of less than total absurdity.

We are drawn to the reluctant conclusion that social context is vital, even in computer programming. Perhaps I should have instead done that half course in sociology with the lady with the lisp? But then, of course, I wouldn’t be writing this text, would I?10

My next post will move through counting in English to English arithmetic. Storage in memory may intrude just a bit.

My 2c, Dr Jo.

1

If I’m compelled to be utterly honest, this characterisation is more than a little cruel. The sociologist I avoided turned out to have a heart of gold, and was a prolific champion of education and literacy, well into her 90’s. She acquired her PhD at the age of 40 in a male-dominated academic environment. She was also a potent anti-apartheid activist. Almost all of our initial impressions and stereotypes turn out to be wrong. It turned out she wasn’t even German either. Young people can be so judgy.

2

Other systems may pop up unexpectedly, for example a 4-cycle if you’re thinking in terms of quarter-hours.

3

Apart from different suffixes for “length of rope or an armspan for measuring” (-raanga), and “turn of a rope when hitched to something'“ (-moora’/-mooru’). Don’t even talk about taro and coconuts.

4

Or that, at best, it’s all just an excuse for vague similes.

5

You might ask the The Glenn A. O. Lean Ethnomathematics Research Centre (GLEM) at the University of Goroka, but they seem to be offline at present. Perhaps you’ll need to visit in person.

6

A pure two cycle system counts along the lines of “one, two, two+one, two+two, two+two+one, …” In his book The End of Error: unum computing, John Gustafson describes his take on the counting system of the Warlpiri from the Northern Territory of Australia. Their numeric concepts are limited to ‘none’, ‘one’, ‘two’, ‘many’, ‘need’ (i.e. negative), ‘junk’ (NaN) and ‘all’ (infinity)—page 93. On the subsequent pages, he then goes on to construct a system using just these concepts that is demonstrably more robust than standard modern floating point arithmetic! This blew my mind, the first time I read it.

7

An interesting study is Mian, from the Ok family of PNG languages. This was pretty much pure “two-cycle”—apart from a 27 body part tally system. Sebastian Fedden describes how the English creole Tok Pisin supplanted it, especially the tally system, which was only ever used for counting units of time, and thus easily displaced.

9

We don’t talk about Raku, the rebranded “Perl 6”. Languages can be as evanescent as bloody butterflies 🦋. Fortran, by the way, is very far from being a dinosaur: not only is it the backbone of a lot of high-performance computing, but if you did some linear algebra today, the chances are very good that under the hood, you used Fortran code (written in the 1980s).

10

Some may see this as advantageous :)

No posts

Read the original on drjo.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.