Monday, August 27, 2007

Some Varieties of Time Machine Worth Having

[Cross-posted to Cliopatria & Digital History Hacks]

I've been invited to join the crack team of bloggers at Cliopatria, so I will be cross-posting there and at Digital History Hacks from time-to-time. I'm excited by the opportunity to develop a series of posts on a topic of general interest to historians, while keeping enough technical content to satisfy my regular readers. So... let's build a time machine!

At some point in the early nineties I copied down a quote by Loren Eiseley in a commonplace book:

A man who has once looked with the archaeological eye will never quite see normally again. He will be wounded by what other men call trifles. It is possible to refine the sense of time until an old shoe in the bunch of grass or a pile of nineteenth-century beer bottles in an abandoned mining town tolls in one's head like a hall clock. This is the price one pays for learning to read time from surfaces other than an illuminated dial. It is the melancholy secret of the artifact, the humanly touched thing. The Night Country 1971:81.


I made a note of the source, but not how I came upon it. I know I wasn't reading Eiseley's work because I used to keep lists of the books that I read. At the time I was studying linguistics and cognitive science, and in the early summer of 1994 I dipped into ecological anthropology. I assume that I came across the quote then. Now I don't really remember the context as clearly as it sounds. I'm making inferences from my old notebooks and from Usenet posts that have been archived online for 15 years. Reading through those old posts reminds me of what I was doing at the time, although I remember being quite a bit cooler than some of my posts make me sound. I wish that that were my own melancholy secret, but at some point in the 1990s I realized that everything that I had ever typed into a computer was going to be saved forever and eventually made available to everyone.

The Eiseley quote stuck with me, and occasionally I would imagine what it would be like to have an 'archaeological eye.' Being given more to science fiction than fantasy, I tended to imagine a mechanism or instrument or device of some sort, rather than a magical object like a crystal ball. Now at this point I should probably stop and reassure you that I know that it may well be impossible to build a time machine in general, and that it is certainly impossible for me to build one. But I think it can sometimes be quite productive to start with something that you know is impossible, and think through some of the implications anyway. As a genre, fiction is ideally suited to this kind of gedankenexperiment; academic monographs less so. Blogs lie somewhere in between. As my fellow Cliopatrian Timothy Burke once wrote, a blog is an ideal "place to publish small writings, odd writings, leftover writings, lazy speculations, half-formed hypotheses." Plus, time machines are a heck of a lot of fun.

When most people think of a time machine, I suspect they probably imagine something like the H. G. Wells version: jump in, set the dial to whenever, hit a button and you are there. This kind of time machine allows (or requires) you to alter the course of events. Sometimes the results are tragic. In the classic Ray Bradbury story "A Sound of Thunder," one of the characters steps on a prehistoric butterfly and changes the future decidedly for the worse. Sometimes the results are comic, as in Connie Willis's re-take of Jerome K. Jerome. A skeptic might point out that if this kind of time travel were ever going to be possible, we'd already be surrounded by people whizzing back from the future to take our fresh water or oxygen, or buy stock in Google, or exhort their younger selves to study harder, or whatever. For historians, the real problem with being able to alter the past is that it would seem to allow for Bill & Ted-style rewriting on a grand scale, and thus make history utterly pointless. The mutability of history, after all, crucially depends on the immutability of the past.

In fact, physicists are split on the possibility of time travel. Some of those who think time travel might be possible suggest that there could be some law of physics that prevents the creation of weird causal loops--you know, the kind where you go back in time to become your own great-great-grandfather or -mother. Stephen Hawking, for example, postulates a "chronology protection conjecture." (For more, see the article by Paul Davies in Scientific American or his subsequent book.) So when I think of an 'archeological eye' I usually imagine something more voyeuristic: the ability to see or hear or in some way measure the events of the past without affecting the outcome.

Years later, let's say around Y2K, I was studying history. Reading Carlo Ginzburg's essay "Clues" reminded me of the Eiseley quote once again. Wouldn't it be cool to write a history based on virtuoso readings of material evidence? (Like Ginzburg, I read a lot of Sherlock Holmes as a kid.) Unfortunately, the only thing that I was arguably a virtuoso at reading was books, and even that was a stretch. Fortunately I was also reading the work of New Institutional Economists at the time. My head was full of ideas of information costs and transaction costs. Since it costs something to learn something, we can never know very much. I had about the same chance of learning to read old shoes or nineteenth-century beer bottles as I did of learning to read sheet music: fairly low. Choosing to specialize in reading one kind of material evidence would preclude learning to read an almost infinite number of other kinds of traces.

What to do? The key word is 'specialize'. As with other kinds of work, there is a division of interpretive labor. In order to make use of material trace evidence, you don't necessarily need to be able to read it yourself, you simply need to be able to find someone who can. With the traditional tools of scholarship it would have been very difficult to assemble a synoptic view of other people's reconstructions of the past from physical evidence. The emergence of search engines like Google drastically lowered those information costs, however. If you type interpret "wear marks" into Google, you will find a reference to a 1958 paper in the British Chiropody Journal on using shoe wear marks to diagnose foot troubles. You'll find a white paper on how to use scattered light to assess surface and bulk defects in various materials, a paper on the use-wear of stone tools, and so on. You'll find, in other words, a world of chiropodists, materials scientists, forensic scientists, engineers, archaeologists and thousands of other kinds of specialists busy reconstructing the past from its material traces. These are people in search of usable past. They care about past events because they have consequences in the present, and the only way they can access that past is by looking for its indexical signs. These experts don't always agree with one another; the mutability of history also depends on the fact that learning is costly. But since our environment is comprised entirely of survivals from the past, it is a kind of time machine, constantly transporting everything from some past into the present. It is one kind of time machine that is worth having... even if it does seem to work in one direction only and is remarkably difficult to use. (For more on the idea of the environment as an archive of material traces see my new book The Archive of Place.)

Next time: the archive as time machine.

Tags: | |

Saturday, August 18, 2007

Perpetual Analytics with Compression

Perpetual analytics is the process of comparing each new item of incoming information to the whole collection at the moment that it is received. IBM scientist Jeff Jonas writes, "there is an ocean of historical data and it is raining, which is to say new data keeps being introduced ... Think of [perpetual analytics] like 'directing the rain drops' as they fall into the ocean – placing each drop in the right place and measuring the ripples (i.e., finding relationships and relevance to the historical knowledge). Discovery is made during ingestion and relevant insight is published at that magical moment." Jonas contrasts this approach with the more traditional process of creating isolated, specialized databases to hold different kinds of information. Over time, these databases tend to become 'silos': many interesting things might be discovered if the information within them could be integrated, but the information costs are too high to do so.

The most powerful implementation of this idea (not to mention the most difficult) would be general-purpose mining at the scale of the internet. I'll leave that for Google or IBM. Instead, I'm going to describe a special-purpose system that operates in a very restricted and small domain.

Imagine browsing through a collection of online primary sources that may be relevant for your research. They could be diary entries, historic newspaper articles or parliamentary records. As you navigate to each new page, a set of links appears in the right sidebar, the way that sponsored advertisements appear in Google search results. Instead of being ads, however, these are links to related primary and secondary sources. If you are reading a letter, for example, there may be links in the sidebar to biographies of the author, recipient or people mentioned in the text. There may be links to other letters written by these people, or to other letters written at the same time and place. If some known event is being described, there may be links to historical accounts of that event. And so on. If you click on one of these sidebar links, a new tab opens in your browser with that source displayed in it, and with links to other sources that are related to it. The sidebar provides ambient information that may be useful without distracting you from the task at hand.

This recommendation system has two very useful features: it is generated automatically and it gets smarter as you use it. Here's what is going on behind the scenes. When you browse to a page, the system stores a copy of the text in a database. If it is the first page you've ever looked at, nothing else happens. When you go to the second page, however, it stores a copy of the text, then uses the normalized compression distance (NCD) to determine how similar the two pages are. (For more on the NCD, see my earlier posts.) As you browse to each new page, a copy is added to the database, and the NCD is calculated for that page and every other that one you've already visited. The sidebar displays links to the closest ones already in the database.

As described so far, this system is able to cluster your own reading, always showing you links to the most relevant stuff that you've already seen. In order to be really useful, you can seed the database with source collections that are likely to be relevant but are too large to be read systematically. For example, if you are working in a particular national and temporal context, you might add all of the entries from a dictionary of historical biography. If you are working in a particular place, you might add complete runs of local newspapers. For specific fields you could add runs of scholarly journals. For groups of people you could add correspondence and diaries.

Furthermore, the system scales up powerfully for collaborative research if the database is shared by everyone working on a particular subject. As each person finds something of interest, it immediately becomes available for recommendation to any of the others, depending on what they are looking at. Built on top of a server-backed version of Zotero, this tool provides one path to leveraging the power of collective intelligences.

Tags: | | | | | |

Tuesday, July 31, 2007

Putting It in Your Own Words

When we teach history students how to take notes for research, we usually tell them to take down direct quotes sparingly, and to put things in their own words instead. Many university writing labs provide training in the art of paraphrasing. One concern is that direct quotes lend themselves to witting or unwitting plagiarism, especially if the paper is being written the night before it's due.

I've always found paraphrasing to be an unsatisfactory exercise because it is in direct tension with close reading. You read the original passage carefully, set it to one side, and then write out the ideas in your own words. At that point you're supposed to re-read the original passage and make sure that you captured the essence. Of course you didn't. As Mark Twain once said, "The difference between the almost-right word & the right word is really a large matter -- it's the difference between the lightning-bug and the lightning." [*] If a student came to me with this example, I'd tell them that there are times when you really should quote rather than paraphrase.

In fact, when I'm taking notes, I usually write down a lot of direct quotes. When I go back to them later, I find that the author's exact words serve as much better reminders of his or her work than paraphrases do. And when I write my first draft of anything, I usually have a lot more quotes than I'm going to want to have in the final version. I know that I'm going to re-read and re-write each passage dozens of times, and that all but the best quotes will be squeezed out in the process.

The problem of putting something in your own words is paralleled in machine learning by a problem known as overfitting. Suppose you work on the production line of a company that makes delicious little chocolates with multi-colored candy shells [Cdn|US]. Even though all of the candies taste the same, your company has come to the conclusion that people pay attention to the color ... they have marketing campaigns based on a preference for eating the red ones last, or the ability to customize the color, or whatever. Your job is to look at the candies as they go by and sort them by color, tossing out any that don't match one of the approved shades. (Sometimes the coloring machine malfunctions and you end up with colors that are more appropriate to your competitor.) Now any hacker in this situation is going to build a robot, so you do. As the candies come down the line, the robot tries to sort them and you provide feedback. If you don't provide enough training, the robot might decide that all of the candies are either blue or red. It is right some of the time, but not enough. That is known as underfitting. If you provide it with too much training on a limited set of examples, it might be correct 100 percent of the time for those examples, but at the cost of memorizing too much detail. Suppose you see five candies in a row, and categorize each as blue. To simplify quite a bit, things that we call "blue" have a wavelength around 475 nanometers. Your robot, however, comes up with five very specific rules: IF WAVELENGTH = 460.83429nm THEN COLOR = blue; IF WAVELENGTH = 483.00089nm THEN COLOR = blue; and so on. Once you turn it loose on a new batch of candies, it is going to start malfunctioning, because it learned too much detail about your original set of examples. It doesn't know what to do if the wavelength is 460.84000nm. This is the problem of overfitting. Now there are a lot of sophisticated methods for avoiding these problems if you are forced to model a limited data set. But the best way to avoid them is to use a lot of training data.

Which brings us back to putting things in your own words. The problem that students encounter with note-taking doesn't have as much to do with quoting vs. paraphrasing as you might think. The problem has to do with not looking at enough sources. If you only consult a handful of sources, then direct quoting might lead you to plagiarism, which would be a case of overfitting. If you paraphrase a handful of sources instead, you may avoid plagiarism but your essay isn't going to be any more nuanced. That is going to lead to underfitting. Either way, a model of a small number of sources is bound to be a bad predictor for the sources that you didn't consult. The only way out is to read more... a lot more. (See my earlier post on "The Difference That Makes a Difference.")

Tags: | | | |

Saturday, July 21, 2007

Import-Export Specialists

James Clifford once said in an interview that he "often function[s] as a kind of import-export specialist between the disciplines" [On the Edges of Anthropology, 55]. I think it's a great description of a particular kind of academic work: finding an idea, tool or technique that is well understood in one context and putting it to use in another. It has particular relevance for the practice of public history.

While thinking about ways of enriching historical practice with digital sources and computation, I've had a lot of occasion to draw on programming, machine learning, and statistical linguistics. In part, these choices reflect my own interests and training before I became a historian. More than that, they're pretty obvious places to look for inspiration. In many ways, digital history is still very textual. It highlights the act of reading, most tools are designed to augment reading or serve as surrogates for it, and outputs are almost always textual in turn. This is as it should be. Most historians (myself included) love to read. Academic history will remain a primarily textual discipline for the foreseeable future.

As I've begun to explore the idea of creating devices and environments that convey a more ambient sense of the past, however, I've had to look a bit further afield for my imports, finding many opportunities to learn from people involved in interaction design, robotics, performance and electronic music. These scholars are often disciplinary import-export specialists in their own right. If you have some time this summer to spend hacking history appliances, here are some good starting points.

Interaction design. Try Bill Moggridge's Designing Interactions and Dan Saffer's Designing for Interaction.

Robotics. The behavior-based approach of Rodney Brooks and his colleagues starts with simple but fully functional creatures interacting with the real world. More complicated systems are built by adding layers of control which subsume lower-level functionality. This strategy lends itself to designing robust interactions between people and history appliances, as I will show in detail in a future post. The related Junkbots, Bugbots and Bots on Wheels is a good source of ideas and techniques.

I also really enjoy reading the blog of Ashish Derhgawen, who comes up with some very creative hacks on a fairly limited budget. This summer he's already figured out a way to use his cellphone as a remote door opener, written a program that can play the classic video game Pong by watching the screen with a webcam, and given one of his robots the ability to respond to claps and whistles.

Performance. The best book that I've found so far for hooking up sensors and actuators to your computer is Tom Igoe and Dan O'Sullivan's Physical Computing. Both of the authors are associated with NYU's Interactive Telecommunications Program, and their focus on live events makes their work particularly useful for people who want to design experiences. The fact that they usually teach artists rather than engineers makes for a very readable work. Igoe's physical computing website is also a great resource.

Electronica. I like listening to electronic music, but hadn't learned anything about it until quite recently. What I've read about its history suggests that it is quite common for electronic musicians to spend a fair amount of their time building new instruments and exploring their creative possibilities. The Cycling '74 website has an interesting collection of resources, including videos, interviews and tutorials. The Create Digital Music webzine is also full of useful stuff. For me, electronica is Ultima Thule: so far out there that I have a hard time finding my most trusted landmarks (i.e., good books on the subject). Pinch and Trocco's Analog Days is an exception.

Tags: | | | | |

Thursday, July 19, 2007

History Appliances: Spöka

On a recent trip to Ikea I came across this awesome little dude. They're selling Spöka as "children's lighting," but it was pretty clear to me that it was one hack short of a history appliance. It has a rechargeable battery, so that you can use it without it being plugged in. If you slide off the rubber skin, there is a light-bulb-shaped plastic housing inside.



The designer thoughtfully created a case which can be opened into three parts and reassembled with nothing more than a small screwdriver.



On the top you'll find a simple push button toggle to turn it on and off.



We want to be able to control the light with the computer, however, so I interrupted the power supply by cutting the circuit to the battery and soldering in a pair of wires (the blue ones). I put a bit of heat-shrink tubing over the joints to make them more resilient. I also knotted the wires to provide strain relief where they will emerge from the case.



When the case is reassembled, the wires can be fed out of the top of the hole where the recharging plug goes in.



After you slide the rubber skin back on, you have an LED-lamp that can be controlled by your computer. If you want to wire it up directly, you might use your parallel port, like Eric Wilhelm does for the haunted house controller in Make volume 3. Instead, I incorporated it into my standard history appliance rig, which uses Phidgets controlled by Max/MSP.

For a quick demo project, I created a browser that lets me look through historic newspaper articles about séances from the online Globe and Mail archive. While browsing the stories from a particular time period, Spöka flashes gently in the background, faster if there are a lot of them, slower if not. It provides a nice peripheral feel for the intensity of Spiritualist activity at that point in time.



Tags: | | | |

Monday, July 02, 2007

Search Refinement with Compression

A few days ago I described a way of using Cilibrasi and Vitányi's Normalized Compression Distance (NCD) to automatically cluster bibliographic entries from the online Dictionary of Canadian Biography. A compression algorithm keeps track of redundancies when it is compressing a string. If those redundancies also occur in another string, then the two strings have something in common (i.e., the redundancies). The NCD ranges from 0 (if the two strings are identical) to 1 (if there is absolutely no overlap). Details are in the original article and laid out in one of my earlier posts.

Compression can also be used to automatically refine searches. Suppose you are interested in the explorer Martin Frobisher. If you type "Frobisher" into Yahoo! some of the first few pages of hits are relevant and some are not. Usually you have to wade through the results (or specify more search keywords and hope you don't eliminate something interesting by being too specific.)

An alternate strategy is to enter a broad search keyword (e.g., "Frobisher") and use the NCD to automatically compare the summary that Yahoo! returns for each hit with a "probe" text such as Frobisher's DCB entry. A short Python program to do exactly that is listed here. The search engine results can then be ranked according to increasing NCD from the probe text.

The figure below shows the first 31 of 50 hits for "Frobisher" before and after this search refinement process. I used red font to indicate the irrelevant results. As can be seen, this use of compression and a probe text does a good job of floating the relevant hits to the top of the pile.



Tags: | | | | | |

Wednesday, June 27, 2007

Clustering with Compression

Last spring I posted short piece about Rudi Cilibrasi and Paul Vitányi's use of compression as a universal method for clustering ["Clustering by Compression," IEEE Transactions on Information Theory 51, no. 4 (2005): 1523-45, PDF]. The basic idea is that a compression algorithm makes a string shorter by keeping track of redundancies and eliminating them. Suppose you have two strings, x and y. If there is some overlap between them, then the concatenated and compressed string xy should be smaller than the concatenation of separately compressed strings x and y. (There are more details in my earlier post). Cilibrasi and Vitányi formalized this idea as the Normalized Compression Distance (NCD).

In my earlier post I selected a handful of entries from the Dictionary of Canadian Biography and submitted them to an open source clustering program that Cilibrasi and Vitányi had provided. I chose the people that I did because I already had some idea of how I would group them myself. The results of the automated clustering were very encouraging, but I hadn't had a chance to follow up with compression-based clustering until now.

This time around, I decided to do a more extensive test. I wrote a Python program to randomly select 100 biographies from volume 1 of the DCB and compute the NCD between each pair. A second Python program combed through the output file of the first to automatically create a Graphviz script that plots all connections below a user-specified threshold. The resulting graph is shown below for NCDs < 0.77.



Looking at different parts of the figure in turn, it becomes clear that this is a remarkably powerful technique, especially considering that the code is very simple and the algorithm has no domain-specific knowledge whatsoever. It knows nothing about history or about the English language, and yet it is able to find connections among biographical entries that are meaningful to a human interpreter.

One isolated chain of biographies consists of people active in Hudson Bay in the late seventeenth century.



A second isolated cluster consists mostly of Englishmen who settled on the coasts of Newfoundland in the early 1600s.



The main cluster is composed almost entirely of people based in various parts of Quebec in the seventeenth century. There is an arm of Acadian settlers, some of whom spent time in Quebec.



There is a somewhat puzzling arm that I haven't really figured out.



And there is the main body of the network and a downward projection that are comprised entirely of seventeenth-century Québécois.





One benefit of this method is the incredible speed with which it executes. It took only seconds to calculate 4,950 NCDs; clustering the entire DCB would require computing 35,427,153 distance measures and would take less than a day to run on my inexpensive home computer. I'll save that hack for another time.

Tags: | | | | |