At 23:59:59 on June 30, 2012, UTC did something ordinary clocks never do: it counted to sixty. The correction was not a surprise; the International Earth Rotation and Reference Systems Service had announced it almost six months earlier, in a bulletin dated January 5, 2012. It was the twenty-fifth second inserted into civil time since the practice began in 1972, needed because the Earth's spin is not quite steady enough for atomic clocks to track without occasional adjustment. On Linux servers across the internet, processors began spinning without getting any work done. At Mozilla, an engineer reported that Java applications built on Hadoop and ElasticSearch had stopped working. Airline check-in suffered too: Amadeus's Altea reservation system, which the company says more than a hundred airlines have implemented, went offline for about an hour, and staff at Qantas and Virgin Australia counters checked passengers in by hand.
Nobody had fabricated anything and nobody had lied about the correction. The failure traced to one missing instruction, pinned down within a day by John Stultz, a Linux kernel developer at IBM. The kernel's leap-second code moved the system clock back one second but never called the internal function, clock_was_set(), that tells the kernel's fine-grained timers the clock has jumped. Every CLOCK_REALTIME timer set with an absolute deadline less than a second away believed that moment had already passed and returned at once; programs using such a timer in a loop, a common way to pause briefly, spun the processor. Stultz, crediting two other developers who had spotted it first, traced the omission to a 2007 change that had taken the call out of the leap-second code to prevent a deadlock. That code had failed once since, in a different way. After the leap second at the end of 2008, one developer gathered reports from 31 users of 53 hard crashes at or near midnight, and a patch posted on January 2, 2009 blamed the code for writing a log message while holding a lock the scheduler needed. That patch moved the log message. It did not put back clock_was_set().
One company had already learned a version of this. Some of Google's clustered systems had stopped accepting work, on a small scale, during the leap second of 2005, and during 2008 a group of its engineers built a workaround that came to be known as the leap smear. Rather than tell any server that a second had been inserted at all, Google's internal time servers added a couple of milliseconds to every update over a window before the leap second, so that by midnight its clocks had already absorbed the extra second. Tested on about 10,000 servers, the smear was set to switch on automatically across Google's production machines for the leap second at the end of 2008.
The correction to UTC, the international civil time standard, was accurate and on schedule. Google chose to keep that fact away from its software, handing its servers a clock that never showed a minute of sixty-one seconds. The Linux kernel took the fact as given, and its leap-second code had never been made to survive what the fact says: that a second, once in a great while, is allowed to happen twice.
The call had been missing since 2007. It took one true second to show it.