Wednesday, November 14, 2007

The Demon of Software

Last time we explored the macroscopic properties of software behavior using the thermodynamics of steam engines as a guide. We saw that in thermodynamics there is a mysterious fluid at work called entropy which measures how run down a system is. You can think of entropy as depreciation; it always increases with time and never decreases. As I mentioned previously, the second law of thermodynamics can be expressed in many forms. Another statement goes as:

dS/dt ≥ 0

This is a simple differential equation that states that the entropy S of an isolated system can only remain constant or must increase with time when an irreversible process is performed by the isolated system. All of the other effective theories of physics can also be expressed as differential equations with time as a factor. However, all of these other differential equations have a "=" sign in them, meaning that the interactions can be run equally forwards or backwards in time. For example, a movie of two billiard balls colliding can be run forwards or backwards in time, and you cannot tell which is which. Such a process is called a reversible process because it can be run backwards to return the Universe to its original state, like backing out a bad software install for a website. But if you watch a movie of a cup falling off a table and breaking into a million pieces, you can easily tell which is which. For all the other effective theories of physics, a broken cup can spontaneously jump back onto a table and reassemble itself into a whole cup if you give the broken shards the proper energy kick. This would clearly be a violation of the second law of thermodynamics because the entropy of the cup fragments would spontaneously decrease. The second law of thermodynamics is the only effective theory in physics that has a "≥" sign in it, which many physicists consider to be the arrow of time. With the second law of thermodynamics, you can easily tell the difference between the past and the future because the future will always contain more entropy or disorder.

Energy too was seen to be subject to the second law of thermodynamics and subject to entropy increases as well. We saw that a steam engine can convert the energy in high-temperature steam into useful mechanical energy by dumping a portion of the steam energy into a reservoir at a lower temperature. Because not all of the energy in steam can be converted into useful mechanical energy, steam engines, like all other heat engines, can never be 100% efficient. We also saw that software too tends to increase in entropy whenever maintenance is performed upon it. Software tends to depreciate by accumulating bugs.

The Second Law of Thermodynamics
The second law does not mean that we can never create anything of value; it just puts some severe limits on the process. Whenever we create something of value, like a piece of software or pig iron, the entropy of the entire Universe must always increase. For example, a piece of pig iron can be obtained by heating iron ore with coke in a blast furnace, but if you add up the entropy decrease of the pig iron with the entropy increase of degrading the chemical energy of the coke into heat energy, you will find that the net amount of entropy in the Universe has increased. The same goes for software. It is possible to write perfect bug-free software by degrading chemical energy in a programmer’s brain into heat and allowing the heat to cool off to room temperature. A programmer on a 2400 calorie diet (2400 kcal/day) produces about 100 watts of heat sitting at her desk and about 20 – 30 watts of that heat comes from her brain. So the next time your peers comment that you are as dim-witted as a 40 watt light bulb after a code review, please be sure to take that as a compliment!

Kinetic Theory of Gasses and Statistical Mechanics
This time we will drill down deeper and use another couple of effective theories from physics – the kinetic theory of gasses and statistical mechanics. We will see that both of these effective theories can be related to software source code at the line of code level. Recall that in 1738 Bernoulli proposed that gasses are really composed of a very large number of molecules bouncing around in all directions. Gas pressure in a cylinder was simply the result of a huge number of molecular impacts from individual gas molecules striking the walls of a cylinder, and heat was just a measure of the kinetic energy of the molecules bouncing around in the cylinder. Bernoulli’s kinetic theory of gasses was not well received by the physicists of the day because many physicists in the 18th and 19th centuries did not believe in atoms or molecules. The idea of the atom goes back to the Ancient Greeks, Leucippus, and his student Democritus, about 450 B.C. and was formalized by John Dalton in 1803 when he showed that in chemical reactions the relative weights of elemental chemical reactants was always the same and proportional to integer multiples. But many physicists had problems with thinking of matter being composed of a “void” filled with “atoms” because that meant they had to worry about the forces that kept the atoms together and that repelled atoms apart when objects “touched”. To avoid this issue, many physicists simply considered “atoms” to be purely a mathematical trick used by chemists to do chemistry. This held sway until 1897 when J. J. Thompson successfully isolated electrons in a cathode ray beam by deflecting them with electric and magnetic fields in a vacuum. It seems that physicists don’t have a high level of confidence in things until they can bust them up into smaller things.

Now imagine a container consisting of two compartments. We fill the left compartment with pure oxygen gas molecules (white dots) and the right compartment with pure nitrogen gas molecules (black) dots.

Figure 1

Next, we perforate the divider between the compartments and allow the molecules of oxygen and nitrogen to mingle.

Figure 2

After a period of time, we will find that the two compartments now contain a gas that is a uniform mixture of oxygen and nitrogen molecules.

Figure 3

This is an example of a spatial entropy increase. The reverse process, that of a mixture of oxygen and nitrogen spontaneously separating into one compartment of pure oxygen and another compartment of pure nitrogen is never observed to occur. Such a process would be a violation of the second law of thermodynamics.

In 1859, physicist James Clerk Maxwell took Bernoulli’s idea one step further. He combined Bernoulli’s idea of a gas being composed of a large number of molecules with the new mathematics of statistics. Maxwell reasoned that the molecules in a gas would not all have the same velocities. Instead, there would be a distribution of velocities; some molecules would move very quickly while others would move more slowly, with most molecules having a velocity around some average velocity. Now imagine that the two preceding compartments (see Figure 1) are filled with nitrogen gas, but that this time the left compartment is filled with cold slow-moving nitrogen molecules (white dots), while the right compartment is filled with hot fast-moving nitrogen molecules (black dots). If we again perforate the partition between compartments, as in Figure 2 above, we will observe that the fast-moving hot molecules on the right will mix with and collide with the slow-moving cold molecules on the left and will give up kinetic energy to the slow-moving molecules. Eventually, both containers will be found to be at the same temperature (see Figure 3), but we will always find some molecules moving faster than the average (black dots), and some molecules moving slower than the average (white dots) just as Maxwell had determined. This is called a state of thermal equilibrium and demonstrates a thermal entropy increase. Just as with the previous example, we never observe a gas in thermal equilibrium suddenly dividing itself into hot and cold segments (the gas can go from Figure 2 to Figure 3 but never the reverse). Such a process would also be a violation of the second law of thermodynamics.

In 1867, Maxwell proposed a paradox along these lines known as Maxwell’s Demon. Imagine that we place a small demon at the opening between the two compartments and install a small trap door at this location. We instruct the demon to open the trap door whenever he sees a fast-moving molecule in the left compartment approach the opening to allow the fast-moving molecule to enter the right compartment. Similarly, when he sees a slow-moving molecule from the right compartment approach the opening, he opens the trap door to allow the low-temperature molecule to enter the left compartment. After some period of time, we will find that all of the fast-moving high-temperature molecules are in the right compartment and all of the slow-moving low-temperature molecules are in the left compartment. Thus the left compartment will become colder and the right compartment will become hotter in violation of the second law of thermodynamics (the gas would go from Figure 3 to Figure 2 above). With the aid of such a demon, we could run a heat engine between the two compartments to extract mechanical energy from the right compartment containing the hot gas as we dumped heat into the colder left compartment. This really bothered Maxwell, and he never found a satisfactory solution to his paradox. This paradox also did not help 19th-century physicists become more comfortable with the idea of atoms and molecules.

Beginning in 1866, Ludwig Boltzmann began work to extend Maxwell’s statistical approach. Boltzmann’s goal was to be able to explain all the macroscopic thermodynamic properties of bulk matter in terms of the statistical analysis of microstates. Boltzmann proposed that the molecules in a gas occupied a very large number of possible energy states called microstates, and for any particular energy level of a gas there were a huge number of possible microstates producing the same macroscopic energy. The probability that the gas was in any one particular microstate was assumed to be the same for all microstates. In 1872, Boltzmann was able to relate the thermodynamic concept of entropy to the number of these microstates with the formula:

S = k ln(N)

S = entropy
N = number of microstates
k = Boltzmann’s constant

These ideas laid the foundations of statistical mechanics.

The Physics of Poker
Boltzmann’s logic might be a little hard to follow, so let’s use an example to provide some insight by delving into the physics of poker. For this example, we will bend the formal rules of poker a bit. In this version of poker, you are dealt 5 cards as usual. The normal rank of the poker hands still holds and is listed below. However, in this version of poker, all hands of a similar rank are considered to be equal. Thus a full house consisting of a Q-Q-Q-9-9 is considered to be equal to a full house consisting of a 6-6-6-2-2 and both hands beat any flush. We will think of the rank of a poker hand as a macrostate. For example, we might be dealt 5 cards, J-J-J-3-6, and end up with the macrostate of three of a kind. The particular J-J-J-3-6 that we hold, including the suit of each card, would be considered a microstate. Thus for any particular rank of hand or macrostate, such as three of a kind, we would find a number of microstates. For example, for the macrostate of three of a kind, there are 54,912 possible microstates or hands that constitute the macrostate of three of a kind.

Rank of Poker Hands
Royal Flush - A-K-Q-J-10 all the same suit

Straight Flush - All five cards are of the same suit and in sequence

Four of a Kind - Such as 7-7-7-7

Full House - Three cards of one rank and two cards of another such as K-K-K-4-4

Flush - Five cards of the same suit, but not in sequence

Straight - Five cards in sequence, but not the same suit

Three of a Kind - Such as 5-5-5-7-3

Two Pair - Such as Q-Q-7-7-4

One Pair - Such as Q-Q-3-J-10

Next, we create a table using Boltzmann’s equation to calculate the entropy of each hand. For this example, we set Boltzmann’s constant k = 1, since k is just a “fudge factor” used to get the units of entropy using Boltzmann’s equation to come out to those used by the thermodynamic formulas of entropy.

Thus for three of a kind where N = 54,912 possible microstates or hands:

S = ln(N)
S = ln(54,912) = 10.9134872

HandNumber of Microstates NProbabilityEntropy = LN(N)Information Change = Initial Entropy - Final Entropy
Royal Flush 4 1.54 x 10-06 1.3862944 13.3843291
Straight Flush 40 1.50 x 10-053.6888795 11.0817440
Four of a Kind 624 2.40 x 10-04 6.4361504 8.3344731
Full House 3,744 1.44 x 10-038.2279098 6.5427136
Flush 5,108 2.00 x 10-038.5385632 6.2320602
Straight 10,200 3.90x 10-039.2301430 5.5404805
Three of a Kind 54,912 2.11 x 10-0210.9134872 3.8571363
Two Pairs 123,552 4.75 x 10-0211.7244174 3.0462061
Pair 1,098,240 4.23 x 10-0113.9092195 0.8614040
High Card 1,302,540 5.01 x 10-01 14.0798268 0.6907967
Total Hands 2,598,964 1.0014.7706235 0.0000000


Examine the above table. Note that higher ranked hands have more order, less entropy, and are less probable than the lower ranked hands. For example, a straight flush with all cards the same color, same suit, and in numerical order has an entropy = 3.6889, while a pair with two cards of the same value has an entropy = 13.909. A hand that is a straight flush appears more orderly than a hand that contains only a pair and is certainly less probable. A pair is more probable than a straight flush because there are more microstates that produce the macrostate of a pair (1,098,240) than there are microstates that produce the macrostate of a straight flush (40). In general, probable things have lots of entropy and disorder, while improbable things, like perfectly bug-free software, have little entropy or disorder. In thermodynamics, entropy is a measure of the depreciation of a macroscopic system like how well mixed two gases are, while in statistical mechanics entropy is a measure of the microscopic disorder of a system, like the microscopic mixing of gas molecules. A pure container of oxygen gas will mix with a pure container of nitrogen gas because there are more arrangements or microstates for the mixture of the oxygen and nitrogen molecules than there are arrangements or microstates for one container of pure oxygen and the other of pure nitrogen molecules. In statistical mechanics, a neat room tends to degenerate into a messy room and increase in entropy because there are more ways to mess up a room than there are ways to tidy it up.

In statistical mechanics, the second law of thermodynamics results because systems with lots of entropy and disorder are more probable than systems with little entropy or disorder, so entropy naturally tends to increase with time.

How the Demon Unleashed the Atomic Bomb
For nearly 100 years physicists struggled with Maxwell’s Demon to no avail. Our next stop will be to see how the concept of information in physics helped to solve the problem, and along the way to finding a solution to Maxwell's Demon, physicists dramatically changed the history of the world. But before we go on, as an IT professional you deal with information all day long, but have you ever really stopped to think what information is? As a start let’s begin with a very simplistic definition.

Information – Something you know

and then see how trying to figure out what information really is accidentally led to the development of the atomic bomb.

Like many of today’s college graduates, Albert Einstein could not initially find a job in physics after graduating from college, so he went to work as a clerk in the Swiss Patent Office from 1902 – 1908, with the hope of one day obtaining a position as a professor of physics at a university. In 1905, he published four very significant papers in the Annalen der Physik, one of the most prestigious physics journals of the time, on the photoelectric effect, Brownian motion, special relativity, and the equivalence of matter and energy. At the time, Einstein rightly figured that these papers would be his ticket out of the patent office, but to his dismay, the 1905 papers were nearly totally ignored by the physics community. This changed when Max Planck, one of the most influential physicists of the time, took interest in Einstein’s work. With Max Planck’s favorable remarks, Einstein was invited to lecture at international meetings and rapidly rose in academia. Beginning in 1908, he embarked upon a series of positions at increasingly prestigious institutions, including the University of Zürich, the University of Prague, the Swiss Federal Institute of Technology, and finally the University of Berlin, where he served as director of the Kaiser Wilhelm Institute for Physics from 1913 to 1933.

At the University of Berlin, Einstein taught a young Hungarian student named Leo Szilárd. In 1922 Leo Szilárd earned his Ph.D. in physics from the University of Berlin with a thesis on thermodynamics. Now Leo Szilárd had a great knack for applying physics to practical problems and coming up with inventions. In 1926 Szilárd decided to go into the refrigerator business by coming up with an improved design. The problem with the home refrigerators of the day was that they did not use the compression and expansion of inert gasses like Freon to do the cooling. Instead, they used very poisonous and noxious gasses like ammonia for that purpose, and the seal where the spinning electric motor axle entered the compressor could leak those poisonous gasses into a kitchen. Szilárd had read a newspaper story about a Berlin family who had been killed when the seal in their refrigerator broke and leaked poisonous fumes into their home. Szilárd figured that a refrigerator without moving parts would eliminate the potential for seal failure because there would be no compressor motor at all, and he began to explore practical applications for different refrigeration cycles with no moving parts. But how do you launch such an enterprise? Szilárd figured what could be better than to enlist the support of the most famous physicist in the world, who just also happened to have been a patent clerk too, namely Albert Einstein. From 1926 - 1933 Einstein and Szilárd collaborated on ways to improve home refrigeration technology. The two were eventually granted 45 patents in their names for three different models. The way these refrigerators worked was that you simply heated one side of the device with a flame, electric coil, or even concentrated sunlight, and the other side of the device got cold. There were no moving parts so there was no need for a compressor motor with a possibly leaky seal. You can actually buy such refrigerators today. Just search for "natural gas refrigerators" on the Internet. They are primarily used for RVs that need to keep food cold when no electricity is available.

Figure 5 – Diagram from Szilárd’s Dec 16, 1927 refrigerator with no moving parts. Notice that Szilárd wisely gave Einstein top billing.

In 1927 the Electrolux Vacuum Cleaner Company (http://www.electroluxappliances.com/) bought one of their patents, freeing up Szilard to take up the new branch of physics known as nuclear physics. Szilárd became an instructor and researcher at the University of Berlin. There he published a paper, On the Decrease of Entropy in a Thermodynamic System by the Intervention of Intelligent Beings in 1929.

Figure 6 – In 1929 Szilard published a paper in which he explained that the process of the Demon knowing which side of a cylinder a molecule was in must produce some additional entropy to preserve the second law of thermodynamics.

In Szilárd's 1929 paper, he proposed that using Maxwell’s Demon, you could indeed build of a 100% efficient steam engine in conflict with the second law of thermodynamics. Imagine a cylinder with just one water molecule bouncing around in it (Figure 6a). First, the Demon figures out if the water molecule is in the left half or the right half of the cylinder. If he sees the water molecule in the right half of the cylinder (Figure 6b), he quickly installs a piston connected to a weight via a cord and pulley. As the water molecule bounces off the piston (Figure 6c) and moves the piston to the left, it slowly raises the weight and does some useful work upon it. In the process of moving the piston to the left, the water molecule must lose kinetic energy in keeping with the first law of thermodynamics and slow down to a lower velocity and temperature than the atoms in the surrounding walls of the cylinder. When the piston has finally reached the far left end of the cylinder it is removed from the cylinder in preparation for the next cycle of the engine. The single water molecule then bounces around off the walls of the cylinder (Figure 6a), and in the process picks up additional kinetic energy from the jiggling atoms in the walls of the cylinder as they kick the water molecule back into the cylinder each time it bounces off the cylinder walls. Eventually, the single water molecule will once again be in thermal equilibrium with the jiggling atoms in the walls of the cylinder and will be on average traveling at the same velocity it originally had before it pushed the piston to the left. So this proposed engine takes the ambient high-entropy thermal energy of the cylinder’s surroundings and converts it into the useful low-entropy potential energy of a lifted weight. Notice that the first law of thermodynamics is preserved. The engine does not create energy; it simply converts the high-entropy thermal energy of the random motions of the atoms in the cylinder walls into useful low-entropy potential energy, but that does violate the second law of thermodynamics. Szilárd's solution to this paradox was simple. He proposed that the process of the Demon figuring out if the water molecule was in the left-hand side of the cylinder or the right-hand side of the cylinder must cause the entropy of the Universe to increase. So “knowing” which side of the cylinder the water molecule was in must come with a price; it must cause the entropy of the Universe to increase. Recall that our first attempt at defining information was simply to say that information was “something you know”. More aptly, we should have said that useful information was simply “something you know”.

On July 4, 1934, Leo Szilárd filed the first patent application for the method of producing a nuclear chain reaction that could produce a nuclear explosion. This patent included a description of "neutron induced chain reactions to create explosions", and the concept of critical mass. In 1938 Otto Hahn, Lise Meitner and Otto Frisch in Germany succeeded in demonstrating that U-235 could indeed fission and possibly form the basis for such a chain reaction. At the time of this discovery, Szilárd and Enrico Fermi were both at Columbia University, and they conducted a very simple experiment that showed significant neutron multiplication when uranium fissioned, proving that the chain reaction was indeed possible and could form the basis for nuclear weapons. Based upon this experiment and the fear that Nazi Germany would soon be embarking on a program to develop an atomic bomb, Leo Szilárd drafted a letter to President Franklin D. Roosevelt for his old business partner and fellow refrigerator designer, Albert Einstein, to sign. The letter, dated August 2, 1939, explained that nuclear weapons were indeed possible and warned of similar Nazi work on such weapons and recommended that the United States immediately begin a development program of its own. This famous letter resulted in the creation of the Manhattan Project which led to the atomic bombs that ended World War II six years later, almost to the day, with the surrender of Japan on August 15, 1945.

Figure 7 – Leo Szilárd’s letter to President Franklin D. Roosevelt signed by Albert Einstein on August 2, 1939.

Information and the Solution to Maxwell’s Demon
Finally, in 1953 Leon Brillouin published a paper with a thought experiment explaining that Maxwell’s Demon required some information to tell if a molecule was moving slowly or quickly. Brillouin defined this information as negentropy, or negative entropy, and found that information about the velocities of the oncoming molecules could only be obtained by the demon by bouncing photons off the moving molecules. Bouncing photons off the molecules increased the total entropy of the entire system whenever the demon determined if a molecule was moving slowly or quickly. So Maxwell's Demon was really not a paradox after all since even the Demon could not violate the second law of thermodynamics. Here is the abstract for Leon Brillouin’s famous 1953 paper:

The Negentropy Principle of Information
Abstract
The statistical definition of information is compared with Boltzmann's formula for entropy. The immediate result is that information I corresponds to a negative term in the total entropy S of a system.

S = S0 - I

A generalized second principle states that S must always increase. If an experiment yields an increase ΔI of the information concerning a physical system, it must be paid for by a larger increase ΔS0 in the entropy of the system and its surrounding laboratory. The efficiency ε of the experiment is defined as ε = ΔI/ΔS0 ≤ 1. Moreover, there is a lower limit k ln2 (k, Boltzmann's constant) for the ΔS0 required in an observation. Some specific examples are discussed: length or distance measurements, time measurements, observations under a microscope. In all cases it is found that higher accuracy always means lower efficiency. The information ΔI increases as the logarithm of the accuracy, while ΔS0 goes up faster than the accuracy itself. Exceptional circumstances arise when extremely small distances (of the order of nuclear dimensions) have to be measured, in which case the efficiency drops to exceedingly low values. This stupendous increase in the cost of observation is a new factor that should probably be included in the quantum theory.

Brillouin proposed that information is the elimination of microstates that a system can be found to exist in. From the above analysis, a change in information ΔI is then the difference between the initial and final entropies of a system after a determination about the system has been made.

ΔI = Si - Sf
Si = initial entropy
Sf = final entropy

Going back to our poker example, let’s compute the amount of information you convey when you tell your opponent what hand you hold. When you tell your opponent that you have a straight flush, you eliminate more microstates than when you tell him that you have a pair, so telling him that you have a straight flush conveys more information than telling him you hold a pair. For example, there are a total of 2,598,964 possible poker hands or microstates for a 5 card hand, but only 40 hands or microstates constitute the macrostate of a straight flush.

Strait Flush ΔI = Si – Sf = ln(2,598,964) – ln(40) = 11.082

For a pair we get:

Pair ΔI = Si – Sf = ln(2,598,964) – ln(1,098,240) = 0.8614040

When you tell your opponent that you have a straight flush you deliver 11.082 units of information, while when you tell him that you have a pair you only deliver 0.8614040 units of information. Clearly, when your opponent knows that you have a straight flush, he knows more about your hand than if you tell him that you have a pair.

Application to Software
Software (source code, config files, lookup tables, etc.) exists as a set of bytes. Each byte can be in one of 256 microstates if we allow for all possible ASCII states of a byte. Some programmers might object that we do not use all 256 possible ASCII characters in programs, but I am going for the most general case here. Since there are hundreds of programming languages, all using a variety of different character sets, the approximation of using all 256 possible ASCII characters is not too bad. For example, there actually is a programming language called whitespace that only uses non-displayed characters such as spaces, tabs and newlines for its source code character set. Such programs appear as blank whitespace in a normal editor, adding an extra layer of source code security. More on whitespace is available at:

https://en.wikipedia.org/wiki/Whitespace_(programming_language)

Now consider a program that is M bytes long. M bytes of software can be in N = 256M microstates, which will be a very large number for any M of appreciable size. That means there are 256M versions of a program that is M bytes long. Consider a medium size program of 30,000 bytes. The number of versions or microstates of a 30,000-byte program are:

N = 25630,000 = 158 x 10 72,245

That is a 158 with 72,245 zeroes behind it! Unfortunately, nearly all of these potential programs will just be a meaningless jumble of characters. In this huge mix, you will also find car ads, the wedding invitations for every couple ever married in the western world, and the Gettysburg Address. If we narrow our search down in the mix to just the true programs, we will find all 30,000-byte programs that ever have been written, or ever will be written, in every programming language that ever has been or ever will be devised using the ASCII character set. Your job as a programmer is to find one of the small set of 30,000-byte programs in the mix that performs the task you desire with a decent response time.

Later we will explore the biological aspects of softwarephysics, but to skip ahead and borrow a biological concept now, we can consider the 158 x 1072,245 possible versions of a 30,000-byte program as the DNA Landscape of the program. DNA stores information in 4 nucleotide or base pair sequences abbreviated as A, C, T, and G. Each base pair in a living thing can be considered a biological bit and can exist in one of the four states A, C, T, or G. Note that living things use base 4 arithmetic as opposed to the binary or base 2 arithmetic common to all of today's computing devices. Today's computers use bits that can be in one of two states "1" or "0". A simple bacterium, like E. coli, contains about 4 million base pairs. Thus the DNA Landscape of E. coli is the set of all DNA sequences that can be made with 4 million base pairs.

N = 44,000,000
N = 1.0 x 10 2,408,240

which is a 1 with 2,408,240 zeroes behind it. Just as with the Landscape of our 30,000-byte program, nearly all of these DNA sequences lead to a dead bacterium that does not work.

The entropy of a piece of software can be measured using the same simplified Boltzmann’s equation with k = 1 that we used for poker hands.

S = ln(N)
N = Number of microstates

The entropy of all possible programs of length M bytes is:

S = ln(256M) = M ln(256) = 5.5452 M

For example, the entropy of all possible 30,000-byte programs comes to

S = 5.5452 M = (5.5452) (30,000) = 166,356

Now of the 158 x 1072,245 possible 30,000-byte programs, we know that only an infinitesimal number of them will provide the desired functionality with a decent response time. Since 158 x 1072,245 is such a huge number, we might as well make the simplifying approximation that there is only one of the possible programs that does the job. Of course, this is not true, but even if there were a billion billion suitable programs in the mix they would pale to insignificance relative to 158 x 1072,245.

So if we assume that there is only one correct version of the program, we can calculate its information content if we remember that the ln(1) = 0:

ΔI of program = Si - Sf = ln(N) - ln(1) = ln(N)

ΔI of program = ln(N) = ln(256M) = M ln(256)
ΔI = 5.5452 M

Since the number of correct versions of a program is always much less than the number of possible versions, our approximation that there is only one correct version of a program is not too bad. For configuration files, frequently there only is one correct version of a file. Notice in the above formula that the information content of a piece of software is proportional to the number of bytes M in the file. This prediction makes intuitive sense.

For our 30,000-byte program the information content comes to:

ΔI in 30,000 bytes = ln(N) = ln(25630,000)
= 30,000 ln(256) = (5.5452) (30,000)

ΔI = 166,356

Thus a 30,000-byte program contains 166,356 units of information, which is a huge amount of information when you compare it to the information in a straight flush that weighs in with only 11.082 units of information. It turns out that a mere 2 bytes of correct software convey 11.090 units of information, which is about that of a straight flush. That means the odds of getting two bytes of software correct by sheer chance are about the same as drawing a straight flush in poker! This is why programming is so difficult.

When you write a program or other piece of software, you are creating information by eliminating microstates that a file could exist in. Essentially you eliminate all of the programs that do not work, like a sculptor who creates a statue from a block of marble by removing the marble that is not part of the statue. The number of possible microstates of a file that constitute the macrostate of being a "correct" version of a program is quite small relative to the vast number of "buggy" versions or microstates that do not work, and consequently has a very low level of entropy and correspondingly a very large amount of information. According to the second law of thermodynamics, software will naturally seek a state of maximum disorder (entropy) whenever you work on software because there are many more “buggy” versions of a file than there are “correct” versions of a file. Large chunks of software will always contain a small number of residual bugs no matter how much testing is performed. Continued maintenance of software tends to add functionality and more residual bugs. Either the second law of thermodynamics directly applies to software, or we have created a very good computer simulation of the second law of thermodynamics - the effects are the same in either case.

An Alternative Concept of Information
Before proceeding, I must mention that in this discussion we are exclusively dealing with Leon Brillouin’s formulation for the concept of information. Unfortunately, there are now several other concepts of information and entropy floating around in science and engineering, and this can cause a great deal of confusion. For example, people who use Information Theory to analyze the flow of information over transmission networks use a slightly different approach to information, and this approach also uses the terms of information and entropy, but in a different manner than we have discussed. In fact, in Information Theory, people actually equate entropy to the amount of useful information in a message! In Information Theory people calculate the entropy, or information content of a message, by mathematically determining how much “surprise” there is in a message. For example, in Information Theory, if I transmit a binary message consisting only of 1s or only of 0s, I transmit no useful information because the person on the receiving end only sees a string of 1s or a string of 0s, and there is no “surprise” in the message. For example, the messages “1111111111” or “0000000000” are both equally boring and predictable, with no real “surprise” or information content at all. Consequently, the entropy, or information content, of each bit in these messages is zero, and the total information of all the transmitted bits in the messages is also zero because they are both totally predictable and contain no “surprise”. On the other hand, if I transmit a signal containing an equal number of 1s and 0s, there can be lots of “surprise” in the message because nobody can really tell in advance what the next bit will bring, and each bit in the message then has an entropy, or information content, of one full bit of information. For more on this see Some More Information About Information. This concept of entropy and information content is very useful for people who work with transmission networks and on error detection and correction algorithms for those networks, but it is not very useful for our discussion. For example, suppose you had a 10-bit software configuration file and the only “correct” configuration for your particular installation consisted of 10 1s in a row like this “1111111111”. In Information Theory that configuration file contains no information because it contains no “surprise”. However, in Leon Brillouin’s formulation of information there would be a total of N = 210 possible microstates or configuration files for the 10-bit configuration file, and since the only “correct” version of the configuration file for your installation is “1111111111” there are only N = 1 microstates that meet that condition. Using the formulas above we can now calculate the entropies of our single “correct” 10-bit configuration file and the entropy of all possible 10-bit configuration files as:

Sf = ln(1) = 0

Si = ln(210) = ln (1024) = 6.93147

So using Leon Brillouin’s formulation for the concept of information the Information content of a single “correct” 10-bit configuration file is:

Si - Sf = 6.93147 – 0 = 6.93147

which, if you look at the above table, contains a little more information than drawing a full house in poker without drawing any additional cards and would be even less likely for you to stumble upon by accident than drawing a full house.

So in Information Theory, a very “buggy” 10 MB executable program file would contain just as much information and would require just as many network resources as transmitting a bug-free 10 MB executable program file. Clearly, the Information Theory formulations for the concepts of information and entropy are less useful for IT professionals than are Leon Brillouin’s formulations for the concepts of information and entropy.

The Conservation of Information Does Not Prevent the Destruction of Useful Information
Another consequence of Brillouin’s reformulation of the second law of thermodynamics that the amount of entropy in the Universe must always increase when anything is changed:

dS/dt ≥ 0

implies that the amount of information in the Universe must also decrease in the Universe whenever something is changed:

dI/dt ≤ 0

Whenever you do something, like work on a piece of software, the total amount of disorder in the Universe must increase, the total amount of information must decrease, and the most likely candidate for that to take place is in the software that you are working on! Truly a sobering thought for any programmer. It seems that the Universe is constantly destroying information because that’s how the Universe tells time!

But the idea of destroying information causes some real problems for physicists, and as we shall see, the solution to that problem is that we need to make a distinction between useful information and useless information. Here is the problem that physicists have with destroying information. Recall, that a reversible process is a process that can be run backwards in time to return the Universe back to the state that it had before the process even began as if the process had never even happened in the first place. For example, the collision between two molecules at low energy is a reversible process that can be run backwards in time to return the Universe to its original state because Newton’s laws of motion are reversible. Knowing the position of each molecule at any given time and also its momentum, a combination of its speed, direction, and mass, we can predict where each molecule will go after a collision between the two, and also where each molecule came from before the collision using Newton’s laws of motion. For a reversible process such as this, the information required to return a system back to its initial state cannot be destroyed, no matter how many collisions might occur, in order for it to be classified as a reversible process that is operating under reversible physical laws.

Figure 8– The collision between two molecules at low energy is a reversible process because Newton’s laws of motion are reversible (click to enlarge)

Currently, all of the effective theories of physics, what many people mistakenly now call the “laws” of the Universe, are indeed reversible, except for the second law of thermodynamics, but that is because, as we saw above, the second law is really not a fundamental “law” of the Universe at all. Now in order for a law of the Universe to be reversible, it must conserve information. That means that two different initial microstates cannot evolve into the same microstate at a later time. For example, in the collision between the blue and pink molecules in Figure 8, the blue and pink molecules both begin with some particular position and momentum one second before the collision and end up with different positions and momenta at one second after the collision. In order for the process to be reversible and Newton’s laws of motion to be reversible too, this has to be unique. A different set of identical blue and pink molecules starting out with different positions and momenta one second before the collision could not end up with the same positions and momenta one second after the collision as the first set of blue and pink molecules. If that were to happen, then one second after the collision, we would not be able to tell what the original positions and momenta of the two molecules were one second before the collision since there would now be two possible alternatives, and we would not be able to uniquely reverse the collision. We would not know which set of positions and momenta the blue and pink molecules originally had one second before the collision, and the information required to reverse the collision would be destroyed. And because all of the current effective theories of physics are time reversible in nature that means that information cannot be destroyed. So if someday information were indeed found to be destroyed in an experiment, the very foundations of physics would collapse, and consequently, all of science would collapse as well.

So if information cannot be destroyed, but Leon Brillouin’s reformulation of the second law of thermodynamics does imply that the total amount of information in the Universe must decrease (dS/dt > 0 implies that dI/dt < 0), what is going on? The solution to this problem is that we need to make a distinction between useful information and useless information. Recall that the first law of thermodynamics maintains that energy, like information, also cannot be created nor destroyed by any process. Energy can only be converted from one form of energy into another form of energy by any process. For example, when you drive to work, you convert all of the low entropy chemical energy in gasoline into an equal amount of useless waste heat energy by the time you hit the parking lot of your place of employment, but during the entire process of driving to work, none of the energy in the gasoline is destroyed, it is only converted into an equal amount of waste heat that simply diffuses away into the environment as your car cools down to be in thermal equilibrium with the environment. So why cannot I simply drive home later in the day using the ambient energy found around my parking spot? The reason you cannot do that is that pesky old second law of thermodynamics. You simply cannot turn the useless high-entropy waste heat of the molecules bouncing around near your parked car into useful low-entropy energy to power your car home at night. And the same goes for information. Indeed, the time reversibility of all the current effective theories of physics may maintain that you cannot destroy information, but that does not mean that you cannot change useful information into useless information.

But for all practical purposes from an IT perspective, turning useful information into useless information is essentially the same as destroying information. For example, suppose you take the source code file for a bug-free program and scramble its contents. Theoretically, the scrambling process does not destroy any information because theoretically it can be reversed. But in practical terms, you will be turning a low-entropy file into a useless high-entropy file that only contains useless information. So effectively you will have destroyed all of the useful information in the bug-free source code file. Here is another example. Suppose you are dealt a full house, K-K-K-4-4, but at the last moment a misdeal is declared and your K-K-K-4-4 gets shuffled back into the deck! Now the K-K-K-4-4 still exists as scrambled hidden information in the entropy of the entire deck, and so long as the shuffling process can be reversed, the K-K-K-4-4 can be recovered, and no information is lost, but that does not do much for your winnings. Since all the current laws of physics are reversible, including quantum mechanics, we should never see information being destroyed. In other words, because entropy must always increase and never decreases, the hidden information of entropy cannot be destroyed.

Entropy, Information and Black Holes
In recent years, the idea that information cannot be destroyed has caused quite a battle in physics as outlined in Leonard Susskind’s The Black Hole War – My Battle With Stephen Hawking To Make The World Safe For Quantum Mechanics (2008). It all began in 1972 when Jacob Bekenstein suggested that black holes must have entropy. A black hole is a mass that is so concentrated that its surrounding gravitational field will allow nothing to escape, not even light. At the time, it was thought that a black hole only had three properties: mass, electrical charge, and angular momentum – a measure of its spin. Now imagine two identical black holes. Into the first black hole, we start dropping a large number of card decks fresh from the manufacturer that are in perfect sort order. Into the second black hole, we drop a similar number of shuffled card decks. When we are finished, both black holes will have increased in mass by exactly the same amount and will remain identical because there will be no change to their electrical charge or angular momentum either. The problem is that the second black hole picked up much more entropy by absorbing all those shuffled card decks than the first black hole, which absorbed only fresh card decks in perfect sort order. If the two black holes truly remain identical after absorbing different amounts of entropy, then we would have a violation of the second law of thermodynamics because some entropy obviously disappeared from the Universe. Bekenstein proposed that in order to preserve the second law of thermodynamics, black holes must have a fourth property – entropy. Since the only thing that changes when a black hole absorbs decks of cards, or anything else, is the radius of the black hole’s event horizon, then the entropy of a black hole must be proportional to the area of its event horizon. The event horizon of a black hole essentially defines the size of a black hole. At the heart of a black hole is a singularity, a point-sized pinch in spacetime with infinite density, where all the current laws of physics break down. Surrounding the singularity is a spherical event horizon. The black hole essentially sucks spacetime down into its singularity with increasing speed as you approach the singularity, and the event horizon is simply where spacetime is being sucked down into the black hole at the speed of light. Because nothing can travel faster than the speed of light, nothing can escape from within the event horizon of a black hole because everything within the event horizon is carried along by the spacetime being sucked down into the singularity faster than the speed of light.

In Bekenstein’s model, the event horizon of a black hole is densely packed with bits, or pixels, of information to account for all the information, or entropy, that has fallen past the black hole’s event horizon. Like in a computer, the pixilated bits on the event horizon of a black hole can be in a state of “1”, or “0”. Each pixel is one Planck unit in area, about 10-70 square meters. In Quantum Software we will learn how Max Planck started off quantum mechanics in 1900 with the discovery of Planck’s constant h = 4.136 x 10-15 eV sec. Later, Planck proposed that, rather than using arbitrary units of measure like meters, kilograms, and seconds that were simply thought up by certain human beings, we should use the fundamental constants of nature - c (the speed of light), G (Newton’s gravitational constant), and h (Planck’s constant) to define the basic units of length, mass, and time. When you combine c, G, and h into a formula that produces a unit of length, called the Planck length, it comes out to about 10-35 meters, so a Planck unit of area is the square of a Planck length or about 10-70 square meters. Now a Planck length is a very small distance indeed, about 1025 times smaller than an atom, so a square Planck length is a very small area, which means the data density of black holes is quite large, and in fact, black holes have the maximum data density allowed in the Universe. They would make great disk drives!

In 1974, Stephen Hawking calculated the exact formula for the entropy of a black hole’s event horizon, and also determined that black holes must also have a temperature, and consequently, radiate energy into space. Recall that entropy was originally defined in terms of heat flow. This meant that because black holes are constantly losing energy by radiating Hawking radiation, they must be slowly evaporating and will one day totally disappear in a brilliant flash of light. This would take a very long time, for example, a black hole with a mass equal to that of the Sun would evaporate in about 2 x 1067 years. But what happens to all the information or entropy that black holes absorb over their lifetimes? Does it disappear from the Universe as well? This is exactly what Hawking, and the other relativists, believed – information, or entropy, could be destroyed as black holes evaporated.

This idea was quite repugnant to people like Leonard Susskind, and other string theorists grounded in quantum mechanics. In future postings, we shall see that the general theory of relativity is a very accurate effective theory that makes very accurate predictions for large masses separated by large distances, but does not work very well for small objects separated by small atomic-sized distances. Quantum mechanics, is just the opposite; it works for very small objects separated by very small distances but does not work well for large masses at large distances. Physics has been trying to combine the two into a theory of quantum gravity for nearly 80 years to no avail. Fortunately, the relativists and quantum people work on very different problems in physics and never even have to speak to each other for the most part. These two branches of physics are only forced to confront each other at the extrema of the earliest times following the Big Bang and at the event horizon of black holes.

In The Black Hole War – My Battle With Stephen Hawking To Make The World Safe For Quantum Mechanics (2008), Susskind describes the ensuing 30-year battle he had with Stephen Hawking over this issue. The surprising solution to all this is the Holographic Principle. The Holographic Principle states that any amount of 3-dimensional stuff in our physical Universe, like an entire galaxy, can be described by the pixilated bits of information on a 2-dimensional surface surrounding the 3-dimensional stuff, just like the event horizon of a black hole. It is called the Holographic Principle because, like the holograms that you see in science museums, the 2-dimensional wiggles of an interference pattern generated by a laser beam bouncing off 3-dimensional objects and recorded on a 2-dimensional film, can regenerate a 3-dimensional image of the objects that you can walk around, and which appears just like the original 3-dimensional objects. In the late 1990s, Susskind and other investigators demonstrated with string theory and the Holographic Principle that black holes are covered with a huge number of strings attached to the event horizon which can break off as photons or other fundamental particles. String theory has been an active area of investigation for the past 30 years in physics and contends that the fundamental particles of nature, such as photons, electrons, quarks, and neutrinos are really very small vibrating strings or loops of energy about one Planck length in size. In this model, the pixilated bits on the event horizon of a black hole are really little strings with both ends firmly attached to the event horizon. Every so often, the strings can twist upon themselves and form a complete loop that breaks free of the event horizon to become a photon, or any other fundamental particle, that is outside of the event horizon and which, therefore, can escape from the black hole. As it does so, it carries away the entropy or information it encoded on the black hole event horizon. This is the string theory explanation of Hawking radiation. In fact, the Holographic Principle can be extended to the entire observable Universe, which means that all the 3-dimensional stuff in our Universe can be depicted as a huge number of bits of information on a 2-dimensional surface surrounding our Universe, like a gargantuan quantum computer calculating how to behave. As I mentioned in So You Want To Be A Computer Scientist, the physics of the 20th century has led many physicists and philosophers to envision our physical Universe simply consisting of information, running on a huge quantum computer. Please keep this idea in mind, while reading the remainder of the postings in this blog.

09/25/2016 Update
For a very interesting update on recent work on black holes and information consider taking the Master Class The Black Hole Information Paradox by Samir Mathur at the World Science U at http://www.worldscienceu.com/.

Conclusion
So what happened to Boltzmann and his dream of a statistical mechanical explanation for the observed laws of thermodynamics? Boltzmann had to contend with a great deal of resistance by the physics establishment of the day, particularly from those like Ernst Mach who had adopted an extreme form of logical positivism. This group of physicists held that it was a waste of time to create theories based upon things like molecules that could not be directly observed. After a lifelong battle with depression and many years of scorn from his peers, Boltzmann tragically committed suicide in 1906. On his tombstone is found the equation:

S = k ln(N)

Scientists and engineers have developed a deep respect for the second law of thermodynamics because it is truly "spooky".

Next time we will further explore the nature of information and its role in the Universe, and surprisingly, see how it helped to upend all of 19th-century physics.

Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston

Friday, October 26, 2007

Entropy - the Bane of Programmers

In my last post, we traced the early history of the science of steam engine building as a progression of mysterious fluids. We began with Johann Joachim Becher’s 1667 mysterious phlogiston which was replaced in 1783 by Lavoisier’s mysterious caloric, and we ended with a mention of Rudolph Clausius’ mysterious fluid called entropy in 1850. Today we will carry on with entropy. As an IT professional, you might find all these mysterious fluids rather amusing, but I would like to remind you that you spend all day working with another mysterious fluid we call information. Information flows through our computer systems on a 24x7 basis, and when it stops flowing, we all get into a lot of trouble rather quickly. When I left United Airlines about 5 years ago, we lost about $150/second when www.united.com went down because customers could not book flights. And I bet that some of you IT folks on Wall Street can easily lose $100,000/second without much trouble. So let’s take our mysterious fluids seriously. The people working on steam engines in the 18th and 19th centuries certainly did. By the way, have you ever stopped to wonder what information is? I mean, you work with it all day long – right? Strange as it might sound, we will soon see how the struggles of steam engine designers in the 19th century led to the concept of information in physics, so please be patient and try to continue learning things from our counterparts in the Industrial Revolution.

In 1842, Julius Robert Mayer unknowingly published the first law of thermodynamics in the May issue of Annalen der Chemie und Pharmacie using experimental results done earlier in France. In this paper, Mayer was the first to propose that there was a mechanical equivalent of heat. Imagine two cylinders containing equal amounts of air. One cylinder has a heavy movable piston supported by the pressure of the confined air, while the other cylinder is completely closed-off. Now heat both cylinders. The French found that the cylinder with the movable piston had to be heated more than the closed-off cylinder to raise the temperature of the air in both cylinders by some identical value like 10 0F. Mayer proposed that some of the heat in the cylinder with the movable piston was converted into mechanical work to lift the heavy piston, and that was why it took more heat to raise the temperature of the air in that cylinder by 10 0F. This is what happens in the cylinders of your car when burning gasoline vapors cause the air in the cylinders to expand and push the pistons down during the power stroke. Mayer proposed there was a new mysterious fluid at work that we now call energy, and that it was conserved. That means that energy can change forms, but that it cannot be created nor destroyed. So for the cylinder with the heavy movable piston, chemical energy in coal is released when it is burned and transformed into heat energy; the resulting heat energy then causes the air in the cylinder to expand, which then lifts the heavy piston. The end result is that some of the heat energy is converted into mechanical energy. Now you might think that with Mayer’s findings that, at long last, steam engine designers finally had some idea of what was going on in steam engines! Not quite. Mayer’s idea of a mechanical equivalent of heat was not well received at the time because Mayer was a medical doctor and considered an outsider of little significance by the scientific community of the day. The sudden loss of two of his children in 1848 and the rejection of his ideas by some of the most prestigious physicists of the time led Mayer to an attempted suicide on May 18, 1850, after which Mayer was committed to an insane asylum.

About the same time another outsider, James Prescott Joule, was doing similar experiments. Joule was the manager of a brewery and an amateur scientist on the side. Joule was investigating the possibility of replacing the steam engines in his brewery with the newly invented electric motor. This investigation ended when Joule discovered that it took about five pounds of battery zinc to do the same work as a single pound of coal. But during these experiments, Joule came to the conclusion that there was an equivalence between the heat produced by an electrical current in a wire, like in a toaster, and the work done by an electrical current in a motor. In 1843, he presented a paper to the British Association for the Advancement of Science in which he announced that it took 838 ft-lbs of mechanical work to raise the temperature of a pound of water by 1 0F (1 Btu). In 1845, Joule presented another paper, On the Mechanical Equivalent of Heat, to the same association in which he described his most famous experiment. Joule used a falling weight connected to a paddle-wheel in an insulated bucket of water via a series of ropes and pulleys to stir the water in the bucket. He then measured the temperature rise of the water as the weight, suspended by a rope connected to the paddle-wheel, descended and stirred the water. The temperature rise of the water was then used to calculate the mechanical equivalent of heat that equated BTUs of heat to ft-lbs of mechanical work. But as with Mayer, Joule’s work was entirely ignored by the physicists of the day. You see, Mayer and Joule were both outsiders challenging the accepted caloric theory of heat.

All this changed in 1847 when physicist Hermann Helmholtz published On the Conservation of Force in which he referred to the work of both Mayer and Joule and proposed that heat and mechanical work were both forms of the same conserved force we now call energy. Finally, in 1850, Rudolph Clausius formalized this idea as the first law of thermodynamics in On the Moving Force of Heat and the Laws of Heat which may be Deduced Therefrom.

The modern statement of the first law of thermodynamics reads as:

dU = dQ – dW

This equation simply states that a change in the internal energy (dU) of a closed system, like a steam engine, is equal to the amount of heat flowing into the system (dQ), minus the amount of mechanical work that the system performs (dW). Thus steam engines take in heat energy from burning coal and convert some of the heat energy into mechanical energy, exhausting the rest as waste heat.

Remember how Carnot thought that the efficiency of a steam engine only depended upon the temperature difference between the steam from the boiler and the temperature of the room in which the steam engine was running? Carnot believed that when caloric fell through this temperature difference, it produced useful mechanical work. Clausius reformulated this idea with a new concept he called entropy, described in his second law of thermodynamics. Clausius reasoned that there had to be some difference in the quality of different energies. For example, there is a great deal of heat energy in the air in the room you are currently sitting in. According to the first law of thermodynamics, it would be possible to build an engine that converts the heat energy in the air into useful mechanical energy and exhausts cold air out the back. With such an engine, you could easily build a refrigerator that produced electricity for your home and cooled your food for free at the same time. All you would need to do would be to hook up an electrical generator to the engine, and then run the cold exhaust from the engine into an insulated compartment! Clearly, this is impossible. As Clausius phrased it, "Heat cannot of itself pass from a colder to a hotter body".

The second law of thermodynamics can be expressed in many ways. One of the most useful goes back to Carnot’s original idea of the maximum efficiency of an engine and the temperature differences between the steam and room temperature. It can be framed quantitatively as:

Maximum efficiency of an engine = 1 – Tc/Th

where Tc is the temperature of the cold reservoir into which heat is dissipated and Th is the temperature of the hot reservoir from which heat is obtained. For a steam engine, Th corresponds to the temperature of the steam and Tc to the temperature of the surrounding room, with both temperatures measured in absolute degrees Kelvin. So Carnot’s proposal was correct. As the temperature difference between Tc and Th gets larger, Tc/Th gets smaller and the efficiency of the engine increases and approaches the value “1” or 100%. In this form of the second law, entropy represents the amount of useless energy; energy that cannot be turned into useful mechanical work and remains as useless waste heat.

Later, in 1865, Clausius presented the most infamous version of the second law at the Philosophical Society of Zurich as:

The entropy of the universe tends to a maximum.

What he meant here was that spontaneous changes tend to smooth out differences in temperature, pressure, and density. Hot objects cool off, tires under pressure leak air and the cream in your coffee will stir itself if you are patient enough. A car will spontaneously become a pile of rust, but a pile of rust will never spontaneously become a car. Spontaneous changes cause an increase in entropy and entropy is just a measure of the degree of this smoothing-out process. So as the entropy of the universe constantly increases with each spontaneous change - the universe tends to run downhill with time. Later we will see that entropy is also a measure of the disorder of a system at the molecular level. In a sense, Murphy’s law is just a popular expression of the second law.

The first and second laws of thermodynamics laid the foundation of thermodynamics which describes the bulk properties of matter and energy. The science of thermodynamics finally allowed steam engine designers to understand what was going on in the cylinders of a steam engine by relating the pressures, temperatures, volumes, and energy flows of the steam in the cylinders while the engine was running. This ended the development of steam engines by trial and error, and the technological craft of steam engine building finally matured into a science.

In Softwarephysics, we apply these same basic ideas of thermodynamics to the macroscopic behavior of software at the program level. The macroscopic behavior of a program can be viewed as the functions the program performs, the speed with which those functions are performed, and the stability and reliability of its performance, just as the macroscopic behavior of steam in a steam engine can be defined by pressure, temperature, and volume changes. This is the viewpoint of software from the perspective of IT management and end-users. They don’t really care what is going on inside of software; they are only interested in how software behaves. From this perspective, it is well known that the entropy of software tends to spontaneously increase with time; software tends to run downhill. Whenever we change software, there is a very good chance that things will run amuck; that is why IT has developed such elaborate change management procedures. But even when we don’t change software, we still seem to get into trouble. I began the very first post on softwarephysics with the following paragraph:

Have you ever wondered why your IT job is so difficult? Have you ever noticed that whenever we change software, performance can drastically decline? Have you observed that performance can drastically decline even when we don’t change software; that sometimes applications spontaneously get slow and then spontaneously return to normal response times without any intervention? Have you noticed that 50% of the time we never find a root cause for problems and that we just start bouncing things at random until performance improves? Have you ever wondered why software behaves this way? Is there anything we can do about all this?

My contention is that these observed behaviors of software are due in part to a simulation of the second law of thermodynamics and the natural tendency for the entropy of software to increase with time at the program level. For deeper insights into this phenomenon of software behavior, we will need to drill down to a lower level and turn to another effective theory of physics called statistical mechanics. Statistical mechanics was developed during the last half of the 19th century and allows us to derive the thermodynamic laws of matter and energy, outlined above, by viewing matter as a large collection of molecules in constant random motion.

Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston

Friday, October 19, 2007

Computer Science as a Technological Craft

A few posts back I described how James Watt made improvements to the Newcomen Steam engine that raised the energy efficiency of steam engines from 1% to 3% and in the process sparked the Industrial Revolution. The reason for exploring the history of steam engines from an IT perspective is two-fold. First of all, it will lead us to some very applicable concepts in the form of the second law of thermodynamics and the interplay of entropy (disorder) and information at the coding level; secondly, it is an interesting story of a technological craft developing into a science. A technological craft is a collection of skills and techniques gained through trial and error that is passed down through the generations. A prime example is early metallurgy. People began to smelt copper about 3800 B.C. in Iran. A thousand years later, in 2800 B.C., the Sumerians in Iraq learned how to combine copper and tin into a very hard and durable alloy called bronze - the distinguishing hallmark of civilization. Around 1500 B.C. the Hittites began to work with iron, which was softer than bronze at the onset, but which made possible the discovery in 1000 B.C. that when iron was reheated with charcoal, it formed a very hard alloy we now call steel. Over many thousands of years of trial and error, people learned many useful metallurgical techniques, like how to harden metal by quenching hot metal in water and then to temper the brittle quenched metal by reheating it and allowing the metal to slowly anneal. Over thousands of years, mankind developed many impressive metallurgical skills and practices, but during all this time, nobody really had any idea of what was going on in the metal during all these intricate manipulations. Acquiring technology by trial and error is a very slow process indeed. Today, metallurgy has evolved into a branch of the material sciences and studies metals at the atomic level within crystal lattices. A modern metallurgist might use x-ray diffraction patterns, electron microscopy, neutron bombardment, or a mass spectrometer to figure out what is going on in a new alloy under study.

Although IT and computer science are populated by very intelligent and gifted people, I would like to suggest that computer science is still an emerging technological craft on the threshold of becoming a true science. In the current state of affairs, even software engineering is little more than the formal instructional framework of a guild. It outlines the procedures to follow to develop and maintain software without dealing with the ultimate nature of software itself. It’s like describing the formal procedures to heat treat the blade of a sword without going into the nature of the steel in the blade itself. In the modern world, civil engineers would find it very difficult indeed to design bridges without knowledge of the tensile strength of structural steel. That is where softwarephysics can be of help to software engineers by providing a theory of software behavior, just as physics comes to the aid of civil engineers by offering a model for the behavior of steel under load.

In 1979, when I transitioned from being an exploration geophysicist to become an IT professional, I left a diverse exploration team exploring for oil in the Gulf of Suez to become a programmer in Amoco’s IT department. Exploration teams are a multidisciplinary team consisting of geologists, geophysicists, petrophysicists, geochemists, and paleontologists. Oil companies throw all the science they can muster at trying to figure out what is going on in a prospective basin before they start spending lots of money drilling holes, just as it is prudent to try to understand the internal nature of an alloy before trying to improve it. When I moved into Amoco’s IT department, I came into contact with many talented and intelligent people, but I was dismayed to discover that there was little sharing of ideas from the other sciences like I had found in my old exploration teams. It seemed that computer science was totally isolated from the outside scientific world. I realized that computer science was a very young science at the time and that it was more of a technological craft than a science, but that was nearly 30 years ago! It’s time for computer science to learn from the other sciences! Fortunately, we are beginning to see this in computer science with the Biologically Inspired Computing (BIC) community in the computer science departments of many universities. The BIC community is, at long last, bringing in ideas from biology, physics, and biochemistry into mainstream computer science and ultimately IT.

Now back to steam engines. Watt did not know that he had increased the energy efficiency of steam engines because he had never even heard of the term energy. The concept of energy did not arrive on the scene until 1850 when Rudolph Clausius published the first law of thermodynamics. But Watt did know that his engine used about 1/3 the coal of a standard Newcomen steam engine with the same horsepower. Unfortunately, the science of the day was not of great help to 18th-century steam engine designers. They had to deal with some very poor effective theories at the time. There was a lot of confusion about the nature of heat in those days. In 1667, Johann Joachim Becher published the phlogiston theory of combustion. The phlogiston theory was an effective theory that proposed that flammable substances such as coal contained a mysterious substance called phlogiston. When coal burned, it released phlogiston to the air. The ash that remained behind was the residual "dephlogisticated" form of coal, while the fumes from the burning coal were "phlogisticated air". Air could only hold so much phlogiston before it became saturated, and that was why a candle could be snuffed out by an overturned glass. The phlogiston theory of combustion held sway until it was replaced by the caloric theory of heat by Antoine Lavoisier in 1783. Lavoisier proposed that when coal burned, it really combined with the newly discovered gas found in air called oxygen and released another mysterious substance called caloric. Caloric was the “substance of heat” and always flowed from hot bodies to cold bodies. Naturally, the hot flue gasses from the burning coal expanded as they took up caloric, making hot air balloons possible. Now during all this time, Daniel Bernoulli had proposed an effective theory of heat that is still in force today known as the kinetic theory of heat. In 1738, Bernoulli proposed that gasses are really composed of a very large number of molecules bouncing around in all directions. Gas pressure in a cylinder was simply the result of a huge number of molecular impacts from individual gas molecules against the walls of a cylinder, and heat was just a measure of the kinetic energy of the molecules bouncing around in the cylinder. But the kinetic theory of heat was not held in favor in Watt’s day, and unfortunately for early steam engine designers, none of these theories of heat related heat energy to mechanical energy, so Watt and the other 18th-century steam engine designers were pretty much on their own, without much help from the available science of the day. In the absence of a useful scientific model for energy, steam engine building in the 18th century became a technological craft, just as computer science has become a technological craft lacking a practical model for software behavior.

This all changed in 1824 when Sadi Carnot published a small book entitled Reflections on the Motive Power of Fire, which was the first scientific treatment of steam engines. In this book, Carnot asked lots of questions. Carnot wondered if a lump of coal could do an infinite amount of work. “Is the potential work available from a heat source potentially unbounded?". He also wondered if the material that the steam engine was made of or the fluid used by the engine made any difference. "Can heat engines be in principle improved by replacing the steam by some other working fluid or gas?". And he came to some very powerful conclusions. Carnot proposed that the efficiency of a steam engine only depended upon the difference between the temperature of the steam used by the engine and the temperature of the room in which the steam engine was running. He envisioned a steam engine as a sort of caloric waterfall, with the difference between the temperature of the steam from the boiler and the temperature of the room the steam engine was in being the height of the waterfall. Useful work was accomplished by the steam engine as caloric fell from the high temperature of the steam to the lower temperature of the room, like water falling over a waterfall doing useful work on a paddlewheel. It did not matter what the steam engine was made of or what fluid was used.

"The motive power of heat is independent of the agents employed to realize it; its quantity is fixed solely by the temperatures of the bodies between which is effected, finally, the transfer of caloric."

“In the fall of caloric the motive power evidently increases with the difference of temperature between the warm and cold bodies, but we do not know whether it is proportional to this difference.”

“The production of motive power is then due in steam engines not to actual consumption of the caloric but to its transportation from a warm body to a cold body.”

Later we will see that Carnot’s ideas were really an early expression of the second law of thermodynamics.

Although we no longer have much confidence in the caloric theory of heat, and now favor the kinetic theory of heat in its place, Carnot was able to make a great contribution to the science of steam engines and thermodynamics by creating an abstract model of steam engines that provided a direction for thought. However, Reflections on the Motive Power of Fire created little attention in 1824 and quickly went out of print. Carnot’s ideas were later revived by work done by Lord Kelvin in 1848 and by Rudolph Clausius in 1850 when Clausius published the first and second laws of thermodynamics in On the Moving Force of Heat and the Laws of Heat which may be Deduced Therefrom. In a similar manner, softwarephysics attempts to provide an abstract model for the behavior of software that provides a direction for thought.

So what happened to Sadi Carnot? In 1832, Carnot died in a cholera epidemic at the age of 36, a victim of the miasma theory of disease described in an earlier post. Unfortunately, because of the fear of choleric miasma, many of his writings were also buried along with him, leaving only a few additional surviving scientific writings.

Next time we will continue on to define the first and second laws of thermodynamics and encounter another strange mysterious substance called entropy – the true bane of all programmers.

Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston

Thursday, October 04, 2007

So Why Are There No Softwarephysicists?

I started working on softwarephysics in 1979 and I have struggled with this question for many years. In my last post, I proposed the idea of thinking of both software and money as being virtual substances. This brought to mind an old high school experience of mine from 1969. During the last semester of my senior year, while I and all of my fellow classmates were well established in our well deserved senior slump, I came across one of those teachers who remain with you for the rest of your life. In this very last semester of high school, we all had to take a mandatory course in economics, in which, from the onset, I totally had no interest, having already mentally departed for the University of Illinois to study physics. However, to my surprise, this teacher totally won me over on the very first day of class with an introduction to the course that went something like this:

We have had money in circulation for several thousand years. We have had domestic and international trade for several thousand years. We have had governments collecting taxes, tariffs, and tribute for several thousand years. We have had banks and money lending for several thousand years. Writing and mathematics have been with us for several thousand years too, and they were invented primarily so that we could do accounting, which has also been with us for several thousand years. And all of these things had life or death consequences for the people of the time. Wars were fought over these issues. Kingdoms and civilizations rose and fell over these issues. All of human history was shaped by these issues. So the question is, dear student, how come there were no economists until the 18th century? Everything that modern economists study and work with had been around for several thousand years, and yet there were no economists! Why? British economist Adam Smith is credited with inventing economics when he published The Wealth of Nations in 1776. Why did that take several thousand years? What changed? What changed was a way of thinking brought about by Galileo, Des Cartes, Spinoza, and Newton. As Thomas Paine put it, it was The Age of Reason. The Enlightenment of the 18th century brought on by the Scientific Revolution of the 17th century created a worldview capable of contemplating economic theories. In this worldview, the Universe was, at last, understandable and rational - it followed physical laws and so too could economic activities.

It has been 66 years since the Age of Software began in the spring of 1941 when Konrad Zuse completed his Z3 computer. But in all that time, I have never seen a theory for software behavior, outside of softwarephysics, that depicted software with invariant tangible physical properties and a set of underlying principles or laws that govern those properties. I have never seen an economic theory of software. Yet deep down I am convinced that nearly all IT professionals really do think of software as a virtual substance. At times we curse software, at other times we cajole software, but at all times we are obsessed with software. I started programming in the fall of 1972 when I took the obligatory course in FORTRAN programming that nearly all physics majors took at the time, and ever since it has been the same story no matter where I go or what I do. I have programmed on punched cards, punched tape, magnetic tape and disk drives using no editor, line editors, full-screen editors, and CASE editors. I have used 3rd generation languages, 4th generation languages, CASE languages, compiled languages and interpreted languages and it has always been predictably the same. When I start working on some code, it’s like nothing has really changed in 35 years. It’s all the same problems over and over, day in and day out. In physics, we would say that software is homogeneous, isotropic and time-invariant. That means that software is the same no matter where you go, where you look, or where you find yourself in time. It’s always the same thing over and over.

Such symmetries have profound implications in modern physics. Emmy Noether was a brilliant mathematical genius, who like Einstein before her, fled Nazi Germany in 1933. In 1918, she published Noether’s Theorem which has become a fundamental tenet in theoretical physics. Her theorem states that there is a one-to-one correspondence between the conservation laws and the symmetries of nature. For example, let’s suppose you do a simple high school physics experiment like colliding two billiard balls on a pool table. Now move the whole contraption 100 miles east and do the same exact experiment. You get the same results. That means the results are symmetric under a spatial translation. Noether showed that this translational symmetry implied the law of the conservation of momentum – the bad stuff that happens when you run into a parked car. Now rotate the pool table 1800 and repeat the experiment. Again you get the same result. Symmetry under rotation implies the conservation of angular momentum – why a skater speeds up when she pulls in her arms in a spin. Do the same experiment a month later, and again you will get the same results. The symmetry over time implies the conservation of energy. So the fact that software is homogeneous, isotropic, and time-invariant just screams out for some underlying laws at work. Something has to be going on!

Next time I really will pick up again with our steam engine designers and thermodynamics. I am a little hesitant to get deeper into physics, since there is the chance of losing some of you, given the very low level of popularity of physics with the general public. But I am going to forge ahead with as little math as possible and try to stick to the main ideas of physics from an IT perspective. My old physics department had a plaque on the wall that read “I understand the material; I just can’t do the problems”. I have frequently thought a more fitting plaque would have been “I can do the problems; I just don’t understand the material.” In my opinion, one of the major failings of teaching physics in this country has been an overemphasis on problem-solving. After all, most physics students never end up being professional physicists, but the underlying concepts of physics can be understood by most people and can be used by all on a daily basis to improve their lives.

So where are all the softwarephysicists? Why you’re one of them! You just don’t know it yet.


Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston

Friday, September 28, 2007

Software as a Virtual Substance

Before diving into thermodynamics, let’s map out the road ahead. Remember that the underlying concept of softwarephysics is that the global IT community has unintentionally created a pretty decent computer simulation of the physical Universe that softwarephysics depicts as the Software Universe. We then use this simulation in reverse. Understanding how the physical Universe operates, allows us to better model how software behaves in the Software Universe. We do this by visualizing software as a virtual substance. Now how do you model a virtual substance? It might help to think of another virtual substance that we are all more familiar with – money. Many people devote their entire educations and careers to learning how to manipulate and model the virtual substance we call money. Nobody gets squeamish about the Federal Reserve Board expanding or contracting the money supply, manipulating interest rates, or trying to raise or lower our exchange rate with other currencies. In fact, you can win a Nobel Prize for such efforts. But for the most part, money is simply a collection of bits stored in a network of computers. Is software any less real than money? Now imagine trying to run the modern world economy without the benefit of macro and microeconomic theories to model the ebb and flow of money throughout the world economies. The aim of softwarephysics is to achieve a similar ability to model the behavior of software at differing levels of software architecture by using the physical Universe as a guide.

The Value of Effective Theories
Our current understanding of the physical universe is based upon a collection of effective theories in physics. Recall that an effective theory is an approximation of reality that only holds true over a certain restricted range of conditions and only provides a certain depth of understanding of the problem at hand. Although all effective theories are fundamentally “wrong”, they are still exceedingly useful. All of the technology surrounding you - your PCs, your cell phones, your GPS units, and your air conditioners, vacuum cleaners, cars, and TV sets were all built using the current approximate effective theories of physics. It’s amazing that you can build all these useful gadgets using theories that are all fundamentally “wrong”, but that just highlights the value of effective theories. The crown jewel of physics is the effective theory called QED - Quantum Electrodynamics. QED has predicted the gyromagnetic ratio of the electron, a measure of its intrinsic magnetic field, to 11 decimal places. This prediction has been validated by rigorous experimental data. As Richard Feynman has pointed out, this is like predicting the exact distance between New York and Los Angeles to within the width of a human hair! Yet most physicists believe that someday an even better effective theory will come along to replace QED. I try to carry over this idea of approximate effective theories into my personal and professional lives too, by trying to keep in mind that my own deeply held personal opinions are also just approximations of reality too. This makes it easier to live with, and work with, people of differing viewpoints.

It might seem disheartening that we don’t really understand what is going on in our own Universe and that all we have is a set of very useful approximations of reality. But as Ayn Rand cautioned us, be sure to “check your premises”. What exactly do you mean by reality? Physicists and philosophers have been debating the nature of reality for thousands of years. My personal favorite definition is:

Physical Reality – Something that does not go away even when you stop believing in it.

I find this definition to be flexible enough to keep most people happy, including physicists, theologians, and even most philosophers. When we discuss the implications of quantum mechanics for software, you may be surprised to learn that about 60% of physicists still ascribe to the old Copenhagen interpretation of quantum mechanics (1927) in which absolute reality does not even exist. In the Copenhagen interpretation, there are an infinite number of potential realities. About 30% adhere to a variation of the “Many-Worlds” interpretation of Hugh Everett (1957) which admits an absolute reality, but claims that there are an infinite number of absolute realities spread across an infinite number of parallel universes. And the remaining 10% bank on the really strange interpretations! They all agree on the underlying mathematics of quantum mechanics; they just don’t agree on what the mathematics is trying to say. It’s like trying to figure out what the mathematical formula:

Glass = 0.5

is trying to tell you. Is the glass half empty or half full? So many physicists skirt the whole issue by adopting a positivist approach to the subject. Logical positivism is an enhanced form of empiricism, in which we do not care about how things “really” are. We are only interested in how things are observed to behave, and effective theories are a perfect fit in that regard. In softwarephysics, I am also taking a positivist approach to software. I don’t care what software “really” is, I only care about how software is observed to behave.

Softwarephysics is a Simulated Science
Just as physics is a collection of useful effective theories, softwarephysics is also a matching collection of effective theories. Because softwarephysics is a simulated science, the challenge for softwarephysics is to find the corresponding matching effective theory of physics to apply to software at each level of complexity. At the highest level, we will be using chaos theory which is a theory for nonlinear dynamical systems developed in the 1970s and 1980s. In the physical Universe, one might apply chaos theory to the traffic patterns on the network of expressways found in a large metropolitan area such as Chicago. In a similar fashion, we could apply chaos theory to the complex traffic patterns for a large corporate website residing on 100 production servers - load balancers, firewalls, proxy servers, webservers, WebSphere Application Servers, CICS Gateway servers to mainframes, mailservers, etc. Any IT professional who has ever worked on a large corporate website might indeed guess that chaos theory would provide a fitting description of their daily life, without ever even having heard of softwarephysics! In fact, chaos theory has already been applied to computer networks by many others.

Thermodynamics
If we pull one of the cars out of a traffic jam on one of the expressways and take a look under the hood, we come to the next level in the hierarchy of effective theories - thermodynamics which was developed during the last half of the 19th century. Thermodynamics allows us to understand the macroscopic behaviors of matter. For example, thermodynamics allows us to understand what is going on in the cylinders of a car by relating the pressures, temperatures, volumes, and energy flows of the gases in the cylinders while the engine is running. We can apply these same basic ideas to the macroscopic behavior of software at the program level. The macroscopic behavior of a program can be viewed as the functions the program performs, the speed with which those functions are performed, and the stability and reliability of its performance.

Statistical Mechanics
The next lower level theory in the hierarchy of effective theories is called statistical mechanics. Statistical mechanics was also developed during the last half of the 19th century and allows us to derive the thermodynamic properties of the gases in the cylinders of a car by viewing the gases as a large collection of molecules bouncing around in the cylinders. It also provides us with a definition of information and some very powerful insights into how information operates in the physical Universe. We will see that we can apply these ideas to software at the line of code level by examining the interplay of information and entropy (disorder) at the coding level.

QED and Chemistry
Going still deeper we come to QED (Quantum Electrodynamics) which is the basis for chemistry at the molecular level. QED is an effective theory that reached maturity in 1948 and which is an amalgam of two other effective theories; quantum mechanics (1925) and Einstein’s special theory of relativity (1905). We will use ideas from QED at the line of code level by depicting lines of code as interacting organic molecules, similar to the chemical reactions back at the refinery that made the gasoline for the car in question.

Quantum Mechanics
Deeper still we finally come to quantum mechanics (1925). Quantum mechanics describes the structure and behavior of individual atoms and we will use concepts from quantum mechanics to describe software at the level of individual characters in source code like the carbon and hydrogen atoms found in gasoline molecules.

In summary:

Software ElementPhysical CounterpartEffective Theory
Computer Networks Expressway Traffic Chaos Theory
ProgramsGas in a Cylinder Thermodynamics and Statistical Mechanics
Lines of CodeOrganic MoleculesQED and Chemistry
Source Code CharactersAtomsQuantum Mechanics


When we are finished with all of this, we will have a collection of effective theories for softwarephysics that will allow us to define a self-consistent model for software behavior that can be used for making day-to-day decisions in IT and provide a direction for thought. It will also lead us to the fundamental problem of software and the suggestion that a biological solution is in order.

Next time we will continue on with the saga of steam engine designers and how their struggle led to the development of the first and second laws of thermodynamics, the concept of entropy (disorder) and the discovery of information itself.

Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston

Sunday, September 23, 2007

A Lesson From Steam Engines

Let’s get back to exploring the benefits of applying science to computer science. Since the rebirth of science about 400 years ago, we have had two major economic revolutions; the Industrial Revolution and the Information Revolution. During the Industrial Revolution, mankind began to manipulate large quantities of energy; while during the Information Revolution, mankind began to manipulate large quantities of information. Both have had huge economic impacts. We can date the dawn of the Information Revolution to the spring of 1941 when Konrad Zuse built the Z3 with 2400 telephone relays. Similarly, we can date the dawn of the Industrial Revolution to 1712 when Thomas Newcomen invented the first commercially successful steam engine.

As an IT professional you are a warrior in the Information Revolution. You work with information all day long. You create, maintain, and operate software which processes information in huge quantities. Software is also a form of information, so essentially you get paid to process information with information. Have you ever stopped to wonder what information is? Is information “real” or just something we made up? Is information a tangible part of the physical Universe, or is it just a useful human contrivance like the names we use for the days of the week? Over the past 400 years, the role of information in physics has taken on more and more significance, to the point that many eminent physicists, such as John Wheeler, have proposed that the physical Universe is simply made out of information - “It from Bit”. Over the years, the concept of information has arisen in physics in several effective theories, most notably in thermodynamics and Einstein’s special theory of relativity. Today we will lay the foundations for the concept of information in thermodynamics, and leave Einstein for another time. Now let’s see if we can learn a lesson from the past warriors of the Industrial Revolution.

The early factories of the 18th century were forced to run on water power. This required them to be located in the highlands near fast-moving water, far from the lowland cities where workers and consumers resided and distant from many natural resources required for production. What was needed was a portable source of power. The Newcomen steam engine was the first commercially successful steam engine and consisted of an iron cylinder with a movable piston. Low-pressure steam was sucked into a cylinder by a rising piston. When the piston reached its maximum extent, a cold water spray was shot into the cylinder causing the steam to condense and form a partial vacuum in the cylinder. External atmospheric air pressure forced the piston down during the power stroke. In the 18th-century, steam engines used low-pressure steam and were thus called atmospheric steam engines because the power stroke came from atmospheric air pressure. High-pressure steam boilers in the 18th century were simply too dangerous to use for steam engines. The Newcomen steam engine was used primarily to pump water out of coal mines. It had an efficiency of about 1%, meaning that about 1% of the energy in the coal used to fuel the engine ended up as useful mechanical work, while the remaining 99% ended up as useless waste heat. This did not bother owners of steam engines in the 18th century because they had never even heard of the term energy. The concept of energy did not come into existence until 1850 when Rudolph Clausius published the first law of thermodynamics. However, they did know that the Newcomen steam engine used a lot of coal. This was not a problem if you happened to own a coal mine, but for 18th-century factory owners, the Newcomen steam engine was far too expensive for their needs.

You can see the oldest surviving Newcomen steam engine at the Henry Ford Museum in Dearborn Michigan just outside of Detroit, as well as Thomas Edison’s original Menlo Park Laboratory, which has also been relocated to the adjoining Greenfield Village museum. This engine was built in 1760 and pumped water from an English coal pit until 1834. I had the chance to see this steam engine a few years ago. It was as big as a house and weighed in at a whopping 15 horsepower, about the horsepower of a modern riding lawnmower. You might wonder why anybody would go to the trouble of building such an engine, but you have to compare it to the effort involved in the care and feeding of 15 horses!

Figure 1 – The first commercially successful steam engine was invented by Thomas Newcomen in 1712. The Newcomen steam engine had an efficiency of 1%.

In 1763, James Watt was a handyman at the University of Glasgow building and repairing equipment for the University. One day the Newcomen steam engine at the University broke, and Watt was called upon to fix it. During the course of his repairs, Watt realized that the main cylinder lost a lot of heat through conduction and that the water spray which cooled the entire cylinder below 212 0F required a lot of steam to reheat the cylinder above 212 0F on the next cycle. In 1765, Watt had one of those scientific revelations in which he realized that he could reduce the amount of coal required by a steam engine if he could just keep the main cylinder above 212 0F for the entire cycle. He came up with the idea of using a secondary condensing cylinder cooled by a water jacket to condense the steam instead of using the main cylinder. He also added a steam jacket to the main cylinder to guarantee that it always stayed above 212 0F for the entire cycle. In 1765, Watt conducted a series of experiments on scale model steam engines that proved out his ideas.

Figure 2 – In 1765, James Watt improved the Newcomen steam engine by introducing a condensing cylinder to condense steam during the power stroke and by using a steam jacket to always keep the main cylinder at a temperature higher than the boiling point of water. Watt's improved steam engine had an efficiency of 3%, and on that basis, launched the Industrial Revolution

To learn moree about the Newcomen and Watt steam engines go to:

https://en.wikipedia.org/wiki/Watt_steam_engine

Watt’s steam engine had an efficiency of 3% which may still sound pretty bad, but that meant it only used 1/3 the coal of a Newcomen steam engine with the same horsepower. So the Watt steam engine became an economically viable option for 18th-century factory owners. We will discuss the second law of thermodynamics at a later time. But just for the sake of comparison, the second law allows us to calculate that the maximum efficiency of a steam engine running at a room temperature of 72 0F using 212 0F steam is 21%.

The Industrial Revolution was delayed by more than 50 years because nobody bothered to try to understand what was going on in a Newcomen steam engine. This was overcome by James Watt when he unknowingly applied the scientific method to steam engines. Based upon some empirical evidence gathered while repairing a Newcomen steam engine, he had a moment of inspiration. He then proceeded to deduce the implications of his revelation and came up with the design for a new kind of steam engine. He then tested his design with a series of controlled experiments.

We are now some 60+ years into the Information Revolution, and like our counterparts in the Industrial Revolution, we are still struggling with the inefficiency of creating and operating software. And like our counterparts, we know that we are very inefficient at software, but we really do not have a clue as to how inefficient we may truly be. Softwarephysics proposes that we stop and take a look into our engine compartment.

The Problem with Common Sense
Like the 18th-century engineers struggling with steam engines, IT professionals have developed a common sense approach to software based upon a set of heuristics. In IT, you quickly learn that when coding software it never works the first time. If you are lucky it will work on the 10th try. If it does not work by the 100th try, you need to look for another profession. But common sense is just another effective theory. Recall that an effective theory is an approximation of reality that only holds true over a certain restricted range of conditions and only provides a certain depth of understanding of the problem at hand. Take a ballpoint pen from your desk and one of your shoes. Drop them both from shoulder height and see which one hits the ground first. For nearly 2,000 years, common sense and the teachings of Aristotle held that the shoe will hit the ground first. It was not until the late 16th century that Galileo demonstrated that they hit the ground at the same time. He also discovered that if you doubled the time of a fall, the distance traveled increased by a factor of four (the square of the time). This was one of the first uses of a mathematical model in physics. The purpose of softwarephysics is to go beyond IT common sense and come up with an effective theory of software behavior at a deeper level.

Next time we will continue on with thermodynamics and see how it led to an effective theory of information and how softwarephysics incorporates that theory into a model for software behavior.

Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston

Monday, September 17, 2007

How To Think Like A Scientist

I was just about to tell you about applying science to computer science when I realized I was getting ahead of myself. First I need to define what I mean by science. As with all of softwarephysics, this is my own operational definition. However, I think it is pretty close to the mainstream concept of what science is as held by the majority of the scientific community.

First of all, science is a way of thinking. Science has a methodology to aid in this way of thinking which has been very successful over the past 400 years. The purpose of the scientific method is to formulate theories or models of the Universe. A scientific model is a simplified approximation of reality that allows people to gain insight into the real structure and operation of true reality. Scientists create models to explain observations, predict future observations, and provide direction for thought. The scientific method is a little different than the way most people think in their daily lives, so let’s examine some of the ways people come up with ideas with a little help from our philosophical friends.

There are three main approaches to gaining knowledge:

1. Inspiration/Revelation
These are ideas that just come out of the blue with no apparent source. I find that IT people are very good at this. For example, on a conference call for a website outage, I am frequently surprised at the incredible level of troubleshooting skill of many of the participants. I frequently wonder to myself “Where did that insight come from?” when somebody nails a root cause out of the blue.

Most of the great ideas in science have also come from inspiration/revelation. For example, in 1900 Max Planck had the insight that he could solve the Ultraviolet Catastrophe by assuming that charged particles in the walls of a room could only oscillate with certain fixed or quantized frequencies. The classical electromagnetic theory of the day predicted that the room you are currently sitting in should be bathed in a lethal level of ultraviolet light and x-rays and that the walls of the room should be at a temperature of absolute zero having turned over all of their available energy into zapping you to death. This was clearly evidence of a theory missing the mark by a wide margin! Planck thought that his fixed frequency solution was just a mathematical trick, but in 1905 Einstein had the revelation that maybe this was not just a trick. Maybe light did not always behave as an electromagnetic wave. Maybe light sometimes behaved like a stream of particles we now call photons that only came in fixed or quantized amounts of energy. The fixed energy of the photons would match up with the fixed frequencies of the charged particles in the walls of your room. In 1924, Louis de Broglie had another revelation and suggested that particles, like electrons, might behave like waves too, just as a stream of photons sometimes behaved like an electromagnetic wave. In 1925, Werner Heisenberg and Erwin Schrödinger developed quantum mechanics based upon these insights, and in 1948 the transistors in your PC were invented at Bell Labs based upon quantum mechanics.

The limits of Inspiration/Revelation:
You never know for sure that your idea is correct.

2. Deductive Rationalism
With deductive rationalism, you make a few postulates which usually come from inspiration/revelation and then you deduce additional ideas or truths from them using pure rational thought. Plato and Des Cartes were big fans of deductive rationalism. It goes like this:

If A = B
And B = C
Then A = C

The limits of deductive rationalism:
In 1931, Kurt Gödel proved that no self-consistent mathematical theory could deduce all truths and that no self-consistent mathematical theory could prove that it was always self-consistent (does not contradict itself). So you cannot deduce all truths.

3. Inductive Empiricism
With inductive empiricism, you make a lot of observations and then reverse the deductive rationalism process. Aristotle and John Locke were big fans of inductive empiricism. If I observe that 99.99% of the time that A = C, then I will assume that A is really equal to C, and I will chalk up the .01% discrepancy to observational error. I don’t know anything about B at this point because I have no observations of B’s state. However, if I make some more observations and find that 99.99% of the time that B = C, then I will infer that B is really equal to C, and therefore, that A is really equal to B too.

If A = C 99.99% of the time
And B = C 99.99% of the time
Then B = C, A = C, and A = B

The limits of empirical induction:
The above may all just be coincidences and you have to have good technology in order to make accurate observations. Most Ancient Greek philosophers did not like inductive empiricism because they thought that all physical measurements on Earth were debased and corrupt. They believed in the power of pure uncorrupted rational thought. This was largely due to the poor level of measurement technology they possessed at the time (they had no Wily). But even in the 17th century when Galileo was demonstrating experiments to his patrons that proved, contrary to Aristotle’s teachings, that all bodies fell with the same acceleration, they thought his experimental demonstrations were magic tricks!

People get into trouble when they only use one or two of the above three approaches to knowledge to make decisions. I know that I do. Politicians have frequently been known to not use any of them at all! The power of the scientific method is that it uses all three of the above approaches to knowledge. Like the checks and balances in the U.S. Constitution, this helps to keep you out of trouble.

The Scientific Method
1. Formulate a set of hypotheses based upon inspiration/revelation with a little empirical inductive evidence mixed in.

2. Expand the hypotheses into a self-consistent model or theory by deducing the implications of the hypotheses.

3. Use more empirical induction to test the model or theory by analyzing many documented field observations or performing controlled experiments to see if the model or theory holds up. It helps to have a healthy level of skepticism at this point. As philosopher Karl Popper has pointed out, you cannot prove a theory to be true, you can only prove it to be false. Galileo pointed out that the truth is not afraid of scrutiny, the more you pound on the truth, the more you confirm its validity.

Effective Theories
The next concept that we need to understand is that of effective theories. Physics currently does not have an all-encompassing unifying theory or model. Researchers are looking for a TOE – Theory of Everything in physics, but currently, we do not have one. Instead, we have a series of pragmatic effective theories. An effective theory is an approximation of reality that only works over a certain range of conditions. For example, Newtonian mechanics allowed us to put men on the Moon, but it cannot explain how atoms work or why the clocks on GPS satellites run faster than clocks on Earth. All of the current theories in physics are effective theories that only work over a certain range of conditions. Physics currently comes in three sizes – Small, Medium, and Large

• Small – less than 10-10 meter and tiny masses
Quantum Mechanics – atomic bombs and transistors
• Medium – 19th-Century Classical Physics
Newtonian Mechanics – space shuttle launches
Maxwell’s Electromagnetic Theory – electric motors
Thermodynamics – air conditioners
• Large – greater than 20,000 miles/sec or very massive objects
Einstein’s General Theory of Relativity – cosmology, black holes, and GPS satellites

Since all of the other sciences are built upon a foundation of underlying effective theories in physics, that means that all of science is “wrong”! But knowing that you are “wrong” gives you a huge advantage over people who know that they are “right” because knowing that you are “wrong” allows you to keep an open mind to search for models that are better approximations of reality.

In addition to covering different ranges of conditions, effective theories also come in different levels of depth with more profound effective theories providing deeper levels of insight. For example, Charles’ Law is a very high-level effective theory that states that at a constant pressure, the volume of a gas in a cylinder is proportional to the temperature of the gas. If you double the temperature of a gas in a cylinder having a freely moving piston, its volume will expand and double in size. A more profound effective theory for the same phenomena is called statistical mechanics which views the gas as a large number of molecules bouncing around in the cylinder. When you double the temperature of the gas, you double the energy of the molecules, so they bounce around faster and take up more room. An even deeper effective theory is called quantum mechanics which views the molecules as standing waves in the cylinder.

The goal of softwarephysics is to provide a pragmatic high-level effective theory of software behavior at a level of complexity similar to that of Charles’ Law. Having an effective theory of software behavior is useful because it allows you to make day-to-day IT decisions with more confidence. For example, suppose you learn 30 minutes before your maintenance window goes down that you have a new EJB that must go into production, but that it corrupts 0.5% of a certain new database transaction. A young programmer on your team quickly produces a “fixed” version of the EJB, but he does not have time to regression test it. Do you put the “fixed” EJB into production, or do you go with the one with the known 0.5% bug with the hope that the corrupted database records can be corrected later? As we shall see later, softwarephysics helps in such situations.

The Most Difficult Thing in Science
The final concept of the scientific method is the most difficult for human beings. In science, you are not allowed to believe in things. You are not allowed to accept models or theories without supporting evidence. However, you are allowed to have a level of confidence in models and theories. For example, I do not “believe” in Newtonian mechanics because I know that it is “wrong”, but I do have a high level of confidence that it could launch me into an Earth orbit. I might get blown up on the launch pad, but like all of our astronauts, I would bet my life on Newtonian mechanics getting me into an Earth orbit instead of plunging me into the Sun if I do my calculations properly! Similarly, I have a low level of confidence in the old miasma theory of disease. In the early 19th century, it was thought by the scientific community that diseases were caused by miasma, a substance found in foul smelling air. And there was a lot of empirical evidence to support this model. For example, people who lived near foul smelling 19th-century rivers were more prone to dying of cholera than people who lived further from the rivers. We had death certificate data to prove that empirical fact. If you were running a cesspool cleaning business in the 19th century, you knew that on the first day of work your rookies were likely to get sick and vomit when they were exposed to the miasma from their first cesspool and a few days later they might come down with a fever and die on you! The miasma theory of disease even had predictive power! If you were running a 19th-century cesspool cleaning business in the middle of a cholera epidemic, and you shut down your operation during the epidemic, while your competitors kept theirs open, you would probably enjoy a larger market share when the epidemic subsided. This just highlights the dangers of relying too heavily on the inductive empiricism approach to gaining knowledge.

As a human being, it is hard not to believe in things. I have been married for 32 years and I have two wonderful adult children. And I truly believe in them all! If somebody confronted me with incontrovertible evidence that one of my children had embezzled funds, my first thought would be that there must be some horrible mistake. However, in scientific matters, you are not allowed this luxury.

The Scientific Method and Softwarephysics
So what does all of this have to do with softwarephysics? Softwarephysics is a high-level effective theory of software behavior. It is a simulated science for the simulated Software Universe that we are all immersed in. Let me explain. In the 1970s, I was an exploration geophysicist writing FORTRAN software to simulate geophysical observations for oil companies. When I transitioned into IT in 1979, it seemed like I was trapped in a frantic computer simulation, just like the ones I used to program for oil companies. After a few months in Amoco’s IT department, I had the following inspiration/revelation:

The Equivalence Conjecture of Softwarephysics

Over the past 70 years, through the uncoordinated efforts of over 50 million independently acting programmers to provide the world with a global supply of software, the IT community has accidentally spent more than $10 trillion creating a computer simulation of the physical Universe on a grand scale – the Software Universe.

I soon realized that I could use this simulation in reverse. By understanding how the physical Universe behaved, I could predict how the Software Universe would react to stimuli, and I proceeded to deduce many implications for software behavior based upon this insight. This was a bit of a role reversal; in physics, we use software to simulate the behavior of the Universe, while in softwarephysics we use the Universe to simulate the behavior of software.

The one problem that I have always had with softwarephysics has been with the confirmation of the model via inductive empiricism. How do you produce and analyze large amounts of documented field observations of software behavior or run controlled experiments for a simulated science? “Hey, Boss I would like to run a double-blind experiment where we install software into production, but only half of it goes through UAT testing. The other half comes straight from the programmers as is, and we don’t know which is which in advance”. Unfortunately, I have always had a full-time job without the luxury of graduate students! So I am relying on 30+ years of personal anecdotal observation of software behavior to offer softwarephysics as a working hypothesis.

Next time I will describe why applying science to computer science is a good idea using the challenges faced by steam engine designers in the 18th century as a case study.

Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston

Saturday, September 15, 2007

So You Want To Be A Computer Scientist?

As professional IT people, we are constantly being called upon to innovate. Unfortunately, at most places where I have worked over the past 32 years, that has meant innovating using conventional ideas – a truly difficult thing to do. I know that softwarephysics can be a bit daunting, especially as presented in SoftwarePhysics 101 – The Physics of Cyberspacetime because the course is designed for several audiences – IT people, physicists, and biologists, and none of these folks talk to each other much. So I would like to break down softwarephysics into some smaller chunks that might be easier to absorb from an IT perspective. I work with a large number of my fellow IT people on a daily basis, and I frequently hear that “Why is this happening to me?” sound in their voices at 3:00 AM. This might help.

Let’s begin where it all started in the spring of 1941 when Konrad Zuse built the Z3 with 2400 electromechanical telephone relays. The Z3 was the world’s first full-fledged computer. You don’t hear much about Konrad Zuse because he was working in Germany during World War II. The Z3 had a clock speed of 5.33 Hz and could multiply two very large numbers together in 3 seconds. It used a 22-bit word and had a total memory of 64 words. It only had two registers, but it could read in and store programs via a punched tape. In 1945, while Berlin was being bombed by over 800 bombers each day, Zuse worked on the Z4 and developed Plankalkuel, the first high-level computer language more than 10 years before the appearance of FORTRAN in 1956. Zuse was able to write the world’s first chess program with Plankalkuel. And in 1950 his startup company Zuse-Ingenieurbüro Hopferau began to sell the world’s first commercial computer, the Z4, 10 months before the sale of the first UNIVAC.

Figure 1 – Konrad Zuse with a reconstructed Z3 in 1961 (click to enlarge)


Figure 2 – Block diagram of the Z3 architecture (click to enlarge)


Now in the past 66 years hardware has improved by a factor of about a billion. You can now go to Best Buy with $500 bucks and buy a machine that is approximately a billion times faster than the Z3 with nearly a billion times as much memory. So how much progress have we made on the software side of computer science in this same period of time? How far have we come since Plankalkuel? Now be careful! A billion seconds is 32 years, and I know that some of you have not quite reached that milestone yet. I would estimate that at most we are perhaps 100 – 1,000 times better off at creating, maintaining, and operating software than Zuse was with writing Plankalkuel on punched tape. And I think I am being generous here. So although we have made great strides in software, how come the hardware guys beat us out by a factor of between 1 – 10 million over the past 60 some years? My suggestion is that this was not a fair fight because the hardware guys were cheating - they were using science! Yes, softwarephysics makes the outrageous suggestion that computer scientists try using science! This has already started to happen in academic computer science with the Biologically Inspired Computing community spread across many universities, but it has not yet filtered down much to the commercial IT community.

Next time I would like to discuss why in the world would you possibly want to apply science to computer science? People working on steam engines in the 18th century asked this very same question.

So what happened to Konrad Zuse? Zuse died in 1995 after making many contributions to computing that you use in your job on a daily basis. You can read about his adventures in computing in his own words at:

http://ei.cs.vt.edu/~history/Zuse.html

Being the unsung genius that he was, Zuse published Calculating Space in 1967, in which he proposed that the physical Universe was a giant computer! This crazy idea has recently been adopted and expanded upon by such huge intellects as physicists John Wheeler, Seth Lloyd, David Deutsch and many others now working on quantum computers. In 1687, Newton published his Principia in which he presented the world with Newtonian mechanics. The Newtonian clockwork model of the Universe, which depicted the world as a huge machine relentlessly moving in deterministic paths, dominated Western thought throughout the 18th and 19th centuries. But the rise of quantum mechanics and chaos theory in the 20th century has recently caused many physicists and philosophers to adopt a new model of the Universe which depicts the Universe as a huge quantum computer constantly calculating how to behave.

Comments are welcome at scj333@sbcglobal.net

To see all posts on softwarephysics in reverse order go to:
https://softwarephysics.blogspot.com/

Regards,
Steve Johnston