RSS Amplifier

Poratbo · Jul 21, 2025

Rapid Technical Advancements

0
Sign in to vote or save

Tanj · Poratbo

We all have heard how Moore’s law has changed our world, a change seemingly like no other. For 60 years the number of transistors in a package has doubled on average every 18 months (from 10 to a trillion). The length of time is so great that even if you qualify the measure and change the denominator, for example logic transistors per watt we still averaged a doubling every 20 months.

Moore’s law has been analyzed many times, I am not going to redo that. Instead I want to look at some other examples of rapid change that can help us put this in perspective:

  • the advances in hard disks from 1956 (RAMAC) to 2016 (the zenith of vertical recording) a different 60 year run

  • the change in speed of solving Mixed Integer Linear Programming, from 1988 to 2004 (a tale of supercomputers and algorithms)

  • the rate of communication over fixed lines

  • the rate of gene sequencing

  • the speed of drilling deep wells, like oil wells

It is easy to think of other areas where huge changes have taken place over time but these 5 provide a reasonable set to look for patterns in how change happens.

I thank Babbage for getting me to look at these, with his article on “Huang’s Law”. I wanted to understand how unusual the AI technology acceleration is.

IBM delivered the RAMAC in 1956. It was not the first hard disk drive, and hard disks were not the first commercial magnetic storage (drum storage was widely used in commercial computers of that era), but it was the first commercial magnetic storage using disks. It stored 5 MB (just under 4MB usable as data), weighed around 1,000kg including necessary control circuits, took 600 msec for an average access, and consumed between 10 and 30 kW of power including its controls and air compressor. By 2016 a 5 TB HDD weighed 0.6kg and had an average latency of 5 msec, running at 9W. We can lump weight and power as a single closely related dimension, so an overall figure of merit would be a million times more storage, in 1000 times less weight/power, running 100 times faster. This produces a 10^11 ratio of improvement over 60 years which is also 18 month average doubling.

By 2016 the rate of improvement severely stalled, with at most a 5x improvement in the 9 years since then, a doubling every 48 months

What went right for 60 years, and what went wrong at the end? The rate of change was not constant. In 1973 IBM made a huge change with the introduction of the Winchester system (named for its design goal of 30/30 internal and removable MB configuration). At this point the drive was about 15x larger capacity, 12x faster seek time, and 30x less power and weight than in 1956, a 5400 increase in 17 years. 12 doublings in 17 years, just a bit faster than 18 months. During this time the changes were conversion to medium-scale integrated transistor circuits, physical reduction in platter thickness, greatly improved magnetic oxide coating, sealed and filtered environment with a head “flying” close to the disk, and a smaller mechanical arm with servo-controlled movement instead of fixed stepping. What is striking is that several of those were qualitative changes with huge advantage and obvious changes to the machine. In the 43 years after that there were arguably few technology changes, you would recognize almost everything, just scaled smaller over time. There was a big jump when IBM introduced read-write heads formed directly on silicon chips in the mid-80s, and when the surface media changed to pure metal alloys with vertical magnetization at the end of the century.

HDDs are made up of many different components, and over time every component followed its own trajectory and opportunities to improve. Heads got smaller, the magnets in the servos got stronger, bearings became free of friction and vibration, arms developed dual level articulation, encoding mechanisms approached theoretical limits, overall packages shrank from 35cm platters in the Winchester to 6cm today and separations shrank from a couple of cm to a couple of mm. All of these subassemblies changed using different inventions - there was no masterplan equivalent to Dennard’s Law which provided 30 years of the arc in Moore’s law.

My best guess for why progress was so close to doubling every 18 months is:

  • diverse ecosystem of suppliers with independent opportunities to advance

  • customer appetite for the steady rate of improvement

That also ties in with why the rate of change has slowed so much in the last 10 years - customers have switched heavily to solid state drives. This shrank the customer pool and left mostly one kind of customer interested in HDDs, the cloud cool storage users. These generated much less money for research and were really only interested in increased density (which is slow to change) and less power (also slow to change).

An interesting hypothesis is that SSDs have not just replaced HDD in sales, but also taken on a similar rate of advance. I leave that as an exercise for the readers.

Linear Programming is one of the major algorithms of the 20th century. It helps optimize things under constraints. You may have encountered the Simplex method at university in math or business classes. George Danzig invented Simplex in the 1950s. Over time it was extended to include problems with some parameters that are integers while others are continuous, which became a more general toolkit of Mixed Integer-Linear Programming. MILP was used for everything from the blending of oil sources at refineries to solving the schedules for trains and buses. In the 1990s interior point algorithms were found which could be faster, and quadratic and gradient descent approaches were developed for non-linear functions.

The GOAT for MILP was Robert Bixby who was the lead programmer for the CPLEX product, which was the unrivaled leader in the market. In the early 2000’s CPLEX was bought by IBM, then Bixby, Gu, and Rothberg left and founded Gurobi, who soon took the lead. Around the time of that change Bixby wrote a retrospective of the era from 1988 to 2004, 16 years, in which he reckoned the power of MILP solvers had increased 5.3 million-fold, or roughly 22 doublings in 16 years. That is one doubling every 9 months. For comparison, and the relatively slow growth rate of Moore’s law or of HDD this would have been 32 years of progress.

A Brief History of Linear and Mixed-Integer Programming Computation

Bixby counted a 3,300x improvement in algorithms and a 1,600x improvement in hardware. Hardware improved at Moore's law rate, while the algorithms adapted to use parallel cores and made huge gains elsewhere. Incidentally, all of that era was using 64-bit floating point.

What I see as the reasons for improvement were:

  • if you factor out Moore’s law, the technology “Linear Programming” advanced by itself at the magical 18 month doubling

  • The economic incentives were very strong and there was an international community of practical mathematicians and economists providing diverse improvements. Abstract maybe, but like the HDD there were many component parts improving in parallel.

It is a technical paper but not that hard to read. It has echoes in AI improvments - indeed the huge-scale optimization techniques from that era enabled the gradient descent approach to training AI. MILP and other solvers continue as commercial products but progress has slowed to around 180x for 20 years or one doubling every 34 months. Perhaps the incentives have fallen off, since AI has attracted many of the same talents and optimizers are effective for most commercial use.

There is also a huge spin-off from classic optimization, which is the methods for training AI, that have derived from gradient descent. These have grown very fast both with hardware and algorithm innovation. They are not classic because they accept approximate solutions over low resolution and massively parallel results, but it is still fascinating to see how related innovation flows downhill also, into new fields.

Data has flowed over fixed lines since the telegraph was deployed for practical use. For much of that time the limit on data rate was not in the technology of the line, but in the humans which operated the telegraph or later the mechanical typewriters used to transcribe the signals. I will somewhat arbitrarily begin around the same time as Moore, 1965, when computers started running connections faster than the 75 baud telegraph system. We can also narrow down to long lines 100km or more, excluding simple circuit interconnects, and look for digital data rates excluding analog telephone.
From 1965 up to now, 60 years, we have gone from 75 baud long lines to about 40 terabits per second, looking only at commercial systems. That is also a doubling every 18 months, for 60 years. Fascinating how we keep seeing similar numbers.

And, by the way, data rates will grow at similar rates for a long time to come. Unlike Moore’s law, where Dennard scaling and atomic sizes predict various dead ends, we are at least 100x below the physical limits to signaling over a single line - at least a decade of exhilarating growth still to come.

Signals on a line had, like HDDs, many technical contributions over time. In the era up to the mid-1980s the lines were copper or microwave line-of-sight, and the improvements were riding on Moore’s law silicon. Silicon could make low noise, efficient amplifiers and equalization circuits that enabled signals to be repeated and restored every few miles on land lines or even undersea. In 1984 the Reagan administration broke up the Bell System which released competition in modems attached to the system. There were already high capacity channels for business carried on microwaves, including AT&T competitors like MCI, so there was rapid expansion into megabit rates with T1 (1.5Mbps) through T3 (45Mbps) although these needed special high quality copper lines to link the customer to the nearest exchange.

In the mid 1980s fiber optics became practical. Eventually the OC1 optical channel became standard in the 1990s with a 155Mbps rate. Optics would be the future from that point since fibers offer around 0.2 dB loss per km and hundreds of GHz of bandwidth. They have very little dispersion (change in arrival time by frequency) and nearly constant losses across a wide band, which are all features that allow very long spans between repeaters and high quality signals delivered at huge bandwidth. In the early 1980s the fiber optic amplifier was invented, which is a length of fiber doped with Erbium (those lanthanides - “rare earths” - just turn up useful everywhere) to act as a laser which could amplify the whole bandwidth of the fiber while staying in the optical form inside the fiber. That was a dramatic improvement making the construction of the line almost independent of the signaling systems used at each end. This freed the line and the endpoints to be improved simultaneously by multiple parts of the ecosystem - basic fiber, erbium-doped fiber amplifiers (EDFA), cable construction and power delivery, signal modulation, and multiple wavelengths subdividing the usable spectrum of the fiber.

New fiber chemistry eliminating all traces of water from the pure quartz allows glass fibers to have 5x the bandwidth (but needs new EDFA designs). ZBLAN, a mixture of heavy metal (lanthanum again!) fluorides could make fibers with a theoretical loss of less than 1dB per 100km - enough to cross an ocean without amplification. We cannot manufacture them in useful quantities yet, and one company is pursuing making them in space for zero gravity to get perfect uniform cooling. Then there are hollow fibers of various kinds which are in use over distances up to about 20km, offering lower latencies because the signal is not just low loss but also travelling at the full speed of light (glass fiber signals are around 60% of light speed).

In terms of spectrum use, we can cram up to 6 bits per Hz of bandwidth with the latest modems for long lines. These make individual channels up to 1.6 Tbps and you can fit around 25 to 30 of those on different colors within the passband of the EDFA. This gives us the latest ocean crossing systems with up to 30Tbps per fiber pair and metro systems around a city, within a distance range needing no amplifier, running up to 80Tbps.

An 80 Tbps signal delivered at 100 mW would be about 10,000 photons at the typical 1.55 micron wavelength, and we have detectors with quantum efficiency better than 10%, so the physical limit to reliability is likely in the range of 100 photons or 100x beyond where we are today. We also have some huge industries like AI willing to consume bandwidths far higher than this - we have single GPUs today with an 80Tbps memory throughput, and datacenters planned with millions of such GPUs.

We have both the technology and the market to expect continuing improvement.

Gene sequencing was first possible as a manual lab procedure with the Sanger approach in 1975, using fragments diffusing through electrically biased gels and detected as the reached the end of the gel channel. The process would run at about 0.5 base-pair per hour. This was partially automated in the first Sanger machines in 1977 but not much faster, around 1 to 2 pairs per hour.

Sanger machines got better automated and reached around 40 base pairs in 1985. The method moved from gel plates to capillary reaching 21,000 pairs per hour in 1995. Those Sanger machines were the ones that delivered the first human genome (almost complete and mostly accurate) in 2003.

By 2005 speed had advanced to around 6M base pairs per hour using a new technology of massively parallel replication in a Roche 454 machine using light flashes to indicate which base pair was in use. Variations of this are dominant today, with an alternative approach also working with observing synthesis developed at Illumina (acquired by Sanger). Illumina machines peaked around 833M base pairs per hour in 2015, using batch sizes with mixtures of roughly 20 genome equivalents taking about 3 days to run. In 2025 the product mix has switched to delivering a single genome per machine in around 6 hours, so it is actually a bit slower at 500M base pairs per hour but cheaper and much more widely available, below $500 per genome.

We could look at this either as the rate of advance in batch throughput, which was roughly 1.5 billion times more over the 40 years from 1975 to 2015, or we could look at the elapsed time per single genome as 1 billion times shorter in 50 years. The first number is a doubling every 16 months for 40 years, the second is a doubling every 20 months for 50 years.

Again we see doubling somewhere around the 18 month interval and we see it running for 40 to 50 years, depending on how you choose to measure it. Gene sequencing was supported by Moore’s law - it has been necessary to have a fabulous increase in data processing to reassemble sequences from overlapping fragmentary data - but we cannot explain the advance as simply a reflection of Moore’s law since there was a core grind of new physics, chemistry, and optics necessary to gene sequencing which had nothing to do with silicon. Somehow a completely different set of innovators sustained that magical 18 month doubling pace for entire careers.

Let’s take a look at something even more distanced from silicon, some huge technology change that could not possibly be a proxy for Moore’s law. The effects of improved computing are pervasive on any research and development, but the sheer physicality of drilling many kilometers through rock seemed like a place where innovations would be fundamentally drawn from other sources, so we could see if the 18-month doubling pattern happens when semiconductors are not at the cutting edge.

This fabulous story is well told in Construction Physics (recommended blog!) “What Learning by Doing Looks Like” so I will just summarize here.

Adding synthetic diamond to rock-drill bits began in the 1970s soon after industrial diamond powder became available. In the early 1980s a Hughes style drill would drill around 6 meters per hour and last about 30 hours, for 180 meters of drilling. By the late 1990s the new Polycrystalline Diamond Compact (PDC) bits were drilling a whole 700m well with one bit in 3 days. Cost had fallen from $470 per meter to $70 per meter (not clear if those were inflation adjusted numbers).

There was then something of a pause for more understanding of the drilling physics and chemistry, and a genius Fred Dupriest revolutionized understanding of the limits to drilling. New, more durable drill heads and the insights of how cutting faster under more pressure would defeat the rock instead of wear the drill meant more speed, while the understanding of the process enabled the operators to make more use of telemetry to properly control the drilling. By 2020 a single bit could cut through granite (much harder material than for the study of progress in the 1990s) and get 1000m deep in about 5 hours or 3000m deep in 110 hours.

These numbers have revolutionized drilling wells and are behind the modern explosive growth in horizontal drilling, fracking, and general rebirth of oil and gas fields that were abandoned decades ago since if you can drill ten times as fast for even less money, the economics change. It also helps that a $40 break even points can be acceptable if historic data and modern seismic scans make new drilling predictably successful. The new faster, cheaper drilling is also making geothermal power much more widely economical.

Clearly there was no doubling every 18 months. Factoring in differences in rock and depth and drill lifetime, it looks like the technology got between 100x and 1000x better in 40 years. Many wells that were simply infeasible became feasible. Drill effectiveness was doubling in 48 to 72 months. That is enough to change the world, and it was sustained over 40 years (and still going), but it shows that some aspects of the real world are literally tough to crack, and the grind of progress is slower than Moore.

I can’t present a definitive conclusion, but it is clear there are many bursts of sustained improvement in human activity. Moore’s law has its reflection in some. It was half of MILP advances and maybe 2/3rds of the data rate increases. However, HDDs made use of semiconductors but was otherwise a composite of many improvements. A half of MILP advances were algorithms independent of semiconductors even if they did need semiconductor advances to enable them. This enablement pattern is common. One set of advances can unlock other advances which are not simply derivative. That is happening with AI, too.

The duration of advancement is also interesting. It feels like the 60 years of Moore’s law should be unprecedented - but the advancements in HDD started earlier and also lasted 60 years. Rock drilling was slower but also has been steady for more than 40 years and is not stopped yet. Data transmission has been rapid, run for 60 years by my arbitrary start date, and has decades of headroom in physics, much less limited than logical computations in Moore’s law devices.

I have a suspicion that the frequent recurrence of the roughly 18 month doubling pattern over decades is driven by how fast humans can plan for and absorb change and the production systems can afford to renew their infrastructure. After all, in a sense the physical world has always had the potential to do what we see today, while the decades it took to develop are mostly constrained by individual and societal inspiration and perspiration. If the real world is stubborn, like drills, we get nowhere near 18-month rates of doubling, while if the real world is just waiting for us to figure it out, we may be reaching some human and society limits around 18 months.

I am also struck how many of my examples seem to start in that 1965 to 1975 era and have had the freedom to double ever since. Is this somehow a failure to find earlier history? Quite likely there will be earlier examples I have missed. Perhaps there were long exponential growth phases in chemistry, or jet engines, or the deployment of electricity, areas that seem likely to me. Still, it feels like something in human technological innovation simply got unbound in the 60s and this pattern of seeking continuous improvements became expected and innovators in many fields were inspired to think that way.

I have heard a philosopher say the lasting impact of Hegel was his notion that history displayed an advance of progress for all humanity, an observation which had not seemed obvious until he made the claim. Indeed, many still have a tendency to imagine a golden era in the past, despite most actually readings of history seeming quite the opposite. Hegel was wrong on many facts due to his narrow access to history of the world, but the world would catch up with his vision of broad progress (if not his proposed causes of it) as people all around the world gained improved access to knowledge from everywhere. Perhaps we have been seeing over the last hundred years a similar broad expectation changing the world of technology. Will it continue? We will probably soon be asking AIs about that. And maybe they won’t see a reason to go as slow as 18 months for doubling.

I would be curious to know what other examples my readers know of. It seems that in specific areas sustained, fast advances are reasonably common. Human ingenuity is very powerful. I wonder what AI will do?

I have some article ideas pinned to my drafts board but which ones and when will have no schedule. I will likely publish whenever an idea seems ready to go and likely to be worth your time to read, and I would like to have at least one per month. All articles will be open to all readers.

No posts

Read the original on tanjb.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.