RSS Amplifier

Dr Jo · Aug 23, 2026

Sing a Song of Sixpence

0
Sign in to vote or save

Dr Jo · Dr Jo

Text within this block will maintain its original spacing when published

Sing a song of sixpence,
A pocket full of rye.
Four and twenty blackbirds
Baked in a pie.
When the pie was opened,
The birds began to sing.
Wasn’t that a dainty dish
To set before the king? — English nursery rhyme, c 18th century

The great war between Lilliput and Blefuscu is less about these tiny nations—which Jonathan Swift located somewhere in the south of the Indian Ocean—and more about how human beings are in continual conflict with their great enemy: themselves.

The starting point for this post, at least when we start to speak computerese, is Danny Cohen’s Internet Experiment Note 137, which is freely available on the Web (title as above), although IEEE will naturally try to charge you more than sixpence for it. He introduced the word ‘endian’ into computer programming: the basic idea is that there are two opposing ways we can store and send numbers, starting with the least significant digits first, or the most. Like all good ideas, it was stolen. From 1726—or the six and twentieth year of the eighteenth century, I guess.

Gulliver’s Travels is still well worth a read. Lilliputians are little-endians, and the evil enemy from Blefuscu are big-endians. Take it away, Jonathan:

Besides, our histories of six thousand moons make no mention of any other regions than the two great empires of Lilliput and Blefuscu. Which two mighty powers have, as I was going to tell you, been engaged in a most obstinate war for six-and-thirty moons past. It began upon the following occasion: It is allowed on all hands, that the primitive way of breaking eggs, before we eat them, was upon the larger end; but his present majesty’s grandfather, while he was a boy, going to eat an egg, and breaking it according to the ancient practice, happened to cut one of his fingers. Whereupon the emperor, his father, published an edict, commanding all his subjects, upon great penalties, to break the smaller end of their eggs. The people so highly resented this law, that our histories tell us, there have been six rebellions raised on that account, wherein one emperor lost his life, and another his crown. These civil commotions were constantly fomented by the monarchs of Blefuscu; and when they were quelled, the exiles always fled for refuge to that empire. It is computed, that eleven thousand persons have, at several times, suffered death, rather than submit to break their eggs at the smaller end. Many hundred large volumes have been published upon this controversy, but the books of the Big-endians have been long forbidden, and the whole party rendered incapable, by law, of holding employments.

We’ll however put that great war aside for a moment, as well as its computer equivalent. We’ll start with something simpler …

In our first tutorial, we saw how computer programmers count from zero—except when they don’t. As simple a decision as where to start counting depends heavily on external circumstances: the culture you’re programming in.

Examining some counting systems around the world, we also saw how messy and random things can become. Let’s now count in English. One, two, three … Have you ever wondered why, as we go past ten, we don’t count something like ‘oneteen’, ‘twoteen’, ‘thirteen’, ‘fourteen’, ‘fifteen’, … ?

Where do eleven and twelve even come from? The word ‘eleven’ apparently comes from Old English ‘ęndleofon’, which means ‘one left over’. ‘Twelve’ is two left. Many say this has Germanic roots. Apparently Lithuanian continues the pattern up to 19, and also has male and female gender for numbers up to 9.1

Aww, a few more quirks in various languages. In Hebrew and Classical Arabic, you use the opposite gender when counting 1–9. In French, there’s the transition from seize to dix-sept and the slightly unusual quatre-vingts for 80. In Hindi, longer Sanskrit names for numbers became fused and contracted over the centuries, producing unique names that simply need to be learned for 1–100. Even Mandarin, which comes pretty close to perfection, has two contextual words for the number 2, and a mandatory placeholder zero (零).2

But why ‘four-and-twenty’ blackbirds? Well, even something like ‘thir-teen’ puts the units before the tens. But if this is the old-fashioned way, surely the real question is “Why the turn-around to ‘24’?” Let’s do some simple maths:

124 +
777 

English reads from left to right (D’Oh!) so why not here left to right working are we? The obvious answer is that we’re using Arabic numerals 0 … 9, and when we swiped these, we also swiped the entire system! Arabic is written right-to-left. The carries still work right-to-left, of course.

So reading a number like ‘124’ in Arabic can be considered little-endian, as was “four and twenty”. It seems that when we swiped Arabic arithmetic, we forced it into our left-to-right system, making it big endian.

Actually, that’s a gross over-simplification. Things are far, far more messy. English got its four-and-twenty little endianness from Germanic roots. But the Norman Conquest in 1066 brought in Romance languages, which are big-endian, through and through.

Gulliver’s war between Lilliput and Blefescu is a thinly-disguised mockery of the enmity between England and France, especially when it comes to Protestantism versus Catholicism. But it fits the endian mess rather well, too.

And this is still an over-simplification—of course.3 Even those who counted “four and twenty” mostly used to put the hundreds and thousands first, messing up our simple interpretation. And—naturally—as modern number usage has swept across the globe, the big numbers tend to be spoken upfront.4 Some say there’s a psychological need to do this. I’d guess though that it’s just people being people. It’s messy. Language is always messy. Always.

If we wrote “124” along the lines of 4̬2̬1 then we’d do our addition and the associated carries from left to right—the old English direction. Now let’s look at computer endianness. This is a bit more exacting than simply writing numbers, because it’s about ordering of components. But let’s initially abuse the term ‘endian’, and think simply about how we might store numbers in computers.

Consider a block of computer memory, in which we might sequentially store single-digit integers. To access those integers, we will need some sort of indexing system—and from that last tutorial, we might count starting conveniently at zero.

You can then see there are two ways we might store a number like 124—the little endian way, putting the four in location 0, and so on; and the big-endian way, with the one in location 0, and so on. We have a choice to make.

And Lo! Manufacturers of computers have chosen ways that are indeed either little- or big-endian. Or mixed-up—there are a few wrinkles. The first is that we will commonly be using binary, and talking about bits, bytes and words.

Everyone knows that modern computers generally store numbers in a binary format: as a series of zeroes and ones—‘binary digits’ usually contracted to ‘bits’, as Claude Shannon popularised in 1948. We’re dealing with powers of two, so we count like this:

0 (don’t forget zero), 1, 10, 11, 100, 101, 110, 111, and so on.

You can work out that say five (101) is just 1×22 + 0×21 + 1×20 = 4 + 1. Test yourself by working out 9 in binary, if you want. Yep, 1001. Now ask yourself “If I’m storing a number in memory, does it make more sense to use little- or big-endian storage?”

The choice may seem to be cultural, but there’s more we need to know, to make sense of our choice. We are inexorably dragged into how we format numbers. And, of course, that too carries a powerful cultural legacy. It seems we can’t escape cultural choices wherever we roam.

When we group bits together, we get a byte.5 Traditionally, the number of bits in a byte has varied so much that we often refer to the modern ‘byte’ of eight bits as an ‘octet’, just to be sure. Early systems used to operate printers and punch-card systems often had six-bit bytes; some had nine. And how should we number the bits?

Then we have actual memory storage in ‘words’—is your head spinning yet?—made up of several bytes. And we’re just getting started, because now that we have bytes and words, we need to be able to store a variety of numbers, and work out ways of storing other things like emojis and alphabetic characters as numbers. It all becomes a bit too much, doesn’t it?

But the point Danny Cohen made was that it’s wise to be consistent throughout. And then we have just two sensible choices: align your bits and bytes and words and more little-to-big, or big-to-little. All of them, in the same order. But perhaps we have a way to dodge the confusion?

Perhaps we should just cry ‘Uncle’ (or the relative of your choice) and decide that we’ll deal with the numbers at an abstract level? Ideally, we should be able to write a layer of software that stands between us and the storage. Someone else has to worry about the finer details—once. Everyone can then simply type in the number, and software does the rest. Simple, No?

To make our encapsulation more solid, we also tend to distinguish among various types of numbers. We are all familiar with integers, which may be negative; and if we’ve had any exposure to science or applied mathematics, we should be familiar with exponential notation in numbers like the fantastically large googol (10100), the more moderately sized Avogadro constant (exactly 6.02214076×1023) or the ridiculously small—in SI units—value for the Planck constant (exactly 6.62607015 10-34 J.Hz-1). The only obvious wrinkle here is that in typing these latter values into most computer programs, we often use expressions like 1e100 rather than 10100.

Of course there will still be issues. Why else would I have written this post? To get some context, try the following.

Actually open up the JavaScript console in the web browser you’re using to read this text. This depends on the browser, but F12 will often work. Otherwise, in Firefox, press Ctrl+Shift+k together; in Chrome, try Ctrl+Shift+j (Mac: Cmd+Option+j); Safari and other browsers may force you to fiddle.6 You may need to click on [Console] too. Now type in:

0.1 + 0.2

What?? If you didn’t get back the result “0.30000000000000004”, then that’s an error. This is because internally, floating point numbers are stored pretty universally as a 64-bit number, according to IEEE754. There are rules.

They are carefully thought out. There are rules for how to store numbers in a variety of formats and word sizes, and how to handle quantities like infinity (positive or negative), and ‘not a number’, or ‘NaN’. They appropriately distinguish between positive and negative zero (0 and –0), and tell us the various ways to round numbers, and how to deal with subnormal numbers, and even which operations are best to support.

But however you slice-and-dice and regulate, there are limits to what you can do. There is no absolutely precise way you can store 0.3 as a single, binary, floating point number, and so we end up with the above anomaly. This also tells us that if you want more than about fifteen or sixteen digits of precision, floating point may not be your best friend. There are, of course, ways around this, and in a later post we’ll explore them, but for now, the take-home message is this:

You cannot always ignore the deep layers!

This, in turn, means that you cannot always “abstract away”. Again and again, we’ll see how failure to acknowledge this simple need to look under the hood confutes, bedevils and bebothers unthinking programmers.

Most modern microprocessors are little-endian, notably those from Intel (and similar shoo-ins), as well as ARM, RISC-V and Apple.7 A weird wrinkle here is that ‘Network Byte Order’ is still officially big-endian! So standard transmission of messages on the Internet using TCP/IP headers specifies big-endian order—in opposition to pretty much every modern processor’s view of its internal world. The software then has to address this, every time. All of this jiggling shouldn’t bother you too much, mostly.

But with programming, you may bump into the problem when you least expect it. Rare errors are not necessarily minor, ignorable errors. A tiny stuffup can cause a lot of trouble.

Consider this. In JavaScript, which is the browser default go-to programming language, we can allocate a 4-byte (32 bit) ‘ArrayBuffer’, into which we might store, say a 32-bit integer. No surprises, so far. It just works. It’s also possible to create both 16-bit and 32-bit ‘views’ of the same buffer. If you now pull out the two 16-bit values, these will differ if the underlying endianness does.

So there’s a special DataView object that you can overlay to take care of this embarrassing little quirk. Provided you know it’s there.8

And then we come to circumstances where people mix up endianness, notably with Bitcoin. Bitcoin is messily little-endian, but multiple other protocols like Apache Thrift (an interface definition language & communication protocol) and Simple Ledger Protocol are big-endian. Cue trouble.

Generally speaking, endian errors are so conspicuous that they don’t often cause catastrophic problems. But where your graphic card won’t work on a PowerPC because of an endian problem, well that’s just sad. And when big-endian bugs need to be fixed in a cryptography library, that’s bad.

The mileage you’ll get from the above will vary. If you’re doing serious stuff with numbers, then you’ll soon realise that the floating point problems we’ve hinted at are practically immense. Conversely, you’ll almost certainly never have to worry about the endianness of internet protocols, or the DataView object, unless you’re doing arcane development work. But we still can take away a few lessons. The obvious ones (apart from the ongoing expectation of confusion, wherever we turn) are:

  1. We just need a variety of practical number formats, notably integer, floating point and fixed point. They are very different, and floating point in particular can surprise you.9

  2. Abstraction won’t always save you.

  3. Ultimately, failures of consistency will tend to bite you when you least expect it. It’s wise to know the details of how things work, and those tiny little wrinkles.

There is a much more abstract lesson from all of this. Recall how our ancestors simply lifted up the little-endian Arabic number system into an incongruous environment, and effectively stuffed things up. Okay, we compensated. Big whoop.

The parable—the real take-home—is however this: if we “cut and paste”, we must make damn sure we understand the context from which we are lifting, and the context we’re pasting into.

Because cut-and-paste is almost always a lazy action, this cognitive requirement almost never happens. This is perhaps the major source of cockups in programming.

My next tutorial is likely best skipped by fundamentalists, as it involves the summoning of demons. This is a pity, as an important focus will be sigils in the Perl programming language, written by Larry Wall, who is not just a programmer and linguist, but also an active Christian who has allowed his religion to bleed over into language design. Whether this is a good or a bad thing remains to be seen, of course. Or perhaps it’s just a thing.

My 2c, Dr Jo.

⌘ This symbol is used to indicate posts where I’ve discussed the flagged topic in more detail

First image is from Wikimedia Commons.

1

Lithuanian also has an entirely separate set of cardinal numbers for counting plural-only nouns, and some nouns like ‘door’ are always plural.

2

And then there are financial numerals. Perfectly logical.

3

If you dig into the details, this becomes extremely messy. When Europeans ‘stole’ the Arabic system, they were well aware that the Islamic world had in turn appropriated an Indian system.† If you look at Al-Hawārī’s Essential Commentary* from 1305, it’s clear that Arabic mathematicians synthesised ideas from Greece, the Middle East and India. There were two common notations for numbers: an abjad (jummal) that allocated the 28 letters of the alphabet to 1, 2, 3, …, 9, 10, 20, 30, … 90, 100, 200, 300, …, 900 and 1000; and the Indian system. The abjad was additive, so 500+40+2, respectively (ث)(م) and (ب) — and this was written in a big-endian format, reading from right to left: ثمب. The Indian numerals are extensively documented. (There was also a third, rūmī system. Ibn alBannāʾ wrote a book about using this). For calculation, they used finger reckoning—which stored intermediate results based on finger positions—but also base 60, and (Taraa!) Indian arithmetic. The last was performed on a “dust board”, a flat surface covered in fine sand or dust. This allowed the sort of positional calculations we’re familiar with.
But the important take-home here is that numeric calculation using Indian/Arabic numerals was entirely subservient to the text. It may be a bit of a stretch to impose ‘endianness’ on ancient views of numbers, so I’ve perhaps been a bit wicked here. But in the excellent “commentary on the commentary” by Mahdi Abdeljaouad and Jeffrey Oaks that I’ve linked to above, they explicitly point out (pp 32–33) that:

Arabic is written right-to-left, so when reading a number like “214” in that language one starts from the 4 in the units place. When Europeans translated Arabic texts on arithmetic into Latin in the Middle Ages they preserved the orientation of figures. This is why “214” did not become “412”.

So perhaps not that wicked, after all. Europeans swiped the lot, without too much thought.

*A commentary on a short book by Abū alʿAbbās Aḥmad ibn Muḥammad ibn ʿUthmān alAzdī alMarrākushī, aka “ Ibn alBannāʾ ”. Al-Hawārī’s full name is ʿAbd alʿAzīz ibn ʿAlī ibn Dāwud alHawārī alMiṣrātī. People had more time then—and likely, more respect.
† Ibn alBannāʾ and Al-Hawārī add right-to-left. Going back to the 5th century CE, there is also evidence that the original numbers were spoken in a little-endian form in Sanskrit, right up to the billions. But in the 11th century Principles of Hindu Reckoning by Kūshyār ibn Labbān, it’s clear that addition there proceeds left-to-right!

4

When we represent Arabic in Unicode (e.g. UTF-8 text), the standard prescribes right-to-left reading of Arabic text (naturally) but also asserts that text containing numbers “is actually bidirectional” and explicitly notes that representation of digits is “left to right”. The rules are extraordinarily complex.

5

When ‘byte’ was coined in 1956, Werner Bucholz deliberately put in a ‘y’ so as not to confuse bite and bit.

6

If you’re on a smartphone, you’ll likely need a specific JavaScript app, as your browser will be almost certainly be crippled.

7

Technically RISC-V (and some other systems) supports both little- and big-endian storage.

8

The entirety of WebAssembly is little endian. And DataView defaults to big-endian reads and writes! There were squabbles at the time.

9

We shouldn’t need a variety of formats. Someday I may dig into John Gustafson’s unum format. Oh wait! He now has three different specifications :)

No posts

Read the original on drjo.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.