Time-saving versus work-inducing software

At a glance, office software like Word, PowerPoint or Excel, are great time savers. Nobody would want to go back to the era before Word Processors?

Unfortunately, I believe that this same software bears part of the blame for our long working hours:

  • Word processors entice people to create too many documents. Microsoft Word is the king of corporate busy work. Wherever I have worked, people got busy crafting all sort of useless internal reports or plans. And, of course, reports must be properly formatted with a title page and an index, just in case someone might print it. And updating old documents can be messy: it almost invariably involves formatting bugs. I am tired of having to check that the font is the same throughout the document. Why can’t machines format documents automatically in a consistent manner? Of course, they can and they have been doing it since the seventies (hint: DocBook, LaTeX, web content management systems like blogs).
  • Spreadsheet software is great for prototyping ideas. If I have half an hour to do an analysis, it is hard to beat Excel. There is a catch however: it is difficult to reuse old spreadsheets with new data. Thus, in most organizations, there is a multiplication of spreadsheets. And spreadsheets tend to grow to include many pages, all poorly documented and fragile. Code reuse is possible, but difficult in Excel. Yet there are perfectly good frameworks for data processing such as R. They are orders of magnitude more powerful and less work intensive.
  • PowerPoint is responsible for 90% of the bad business presentations. Have you noticed that Bill Gates frequently give talks without PowerPoint? In fact, he became a much better speaker since he stopped using so many silly slides. But what is worse is that people spend a lot of time on these slides instead of preparing good talks. And remember: not giving a talk is often the best option.

Microsoft is not the sole company to blame. In universities, most assignments and exams are still marked by hand whereas we have had the technology to automate 90% of the marking for years!

Happily, I find that some software really does save labor:

  • Most web content management systems let the author write and publish efficiently. Maintaining this blog is cost-effective: with only a few hours of work every week, I can reach thousands. I spend almost no time on repetitive tasks.
  • Scripting has gotten a lot better in the last 20 years, and it is very useful. I get a lot of my data processing done in Python. My only regret is that so few people learn scripting languages.
  • Obviously, Wikipedia is amazing at saving time.
  • With Doodle, scheduling meetings is an order of magnitude faster than with Microsoft Outlook.
  • Cell phones are work-inducing, obviously. However, I conjecture that tablet-based computing is time-saving. People write shorter comments and emails. They tend to start fewer documents. Users of an iPad will spend more time reading than writing. Isn’t it about time that we take some time off to read instead of producing more than others can consume?

What is the underlying thread? Time-saving software tends to be produced by less civilized people.  Software written by large corporations will probably be work-inducing.

Further reading: Of Lisp Macros and Washing Machines (via Hosh Hsiao) and Conway’s law (via John D. Cook)

Scaling MongoDB

I have been spending much time thinking about a future where document-oriented databases are the default. Though they have their problems, I think that they are far better suited for what most people want to do than relational databases.

MongoDB is one of the best document-oriented database system around: it is mature, scalable, open source and commercially supported.  You can set it up to run Amazon’s cloud in minutes.

A long time ago, the good people at O’Reilly sent me a copy of Scaling MongoDB by Kristina Chodorow, one of the developers of MongoDB. Kristina has a total of four books at O’Reilly.  This one is short, but the writing is nearly perfect. The book looks beautiful too. And no, not all O’Reilly books are that good.

If you are new to MongoDB, you should get a more complete book like MongoDB: The Definitive Guide. Scaling MongoDB is about the specific, but important issues raised by scalability. Having a separate book makes sense. Indeed, Learning the basics of MongoDB is not hard. But figuring it out sharding is a lot harder and deserves a book.

Improve your impact with abundance-based design

People design all the time: new cars, new software, new houses. All design is guided by constraints (cost, time, materials, space) and by objectives (elegance, quality). Constraints are limitations: you only have so much money, so many days… whereas objectives are measures that you seek to either maximize or minimize. In practice, either the constraints or the objectives may dominate. You are either worried about limited ressources, or you seek to maximize the quality of your result.

Our ancestors were probably often forced into scarcity-based design. When your very survival is in question, you build whatever shelter you can in the hours you have left before nightfall. We are probably wired for good scarcity-based design as it is a survival trait.

Any monkey can live in scarcity. However, abundance-based design is crucial if you want to maximize your impact.

Facebook engineers do abundance-based design. They are mainly worried about improving Facebook and pursuing objectives such as usability, but much less worried about time or disk space. Similarly, when I build a model sailboat, costs and time are nearly irrelevant, I mostly care that my boat be pretty and that it handles well. As a researcher, most of my research papers are the result of abundance-based design. It does not matter how long I work on the research projects, as long as the result has impact. Similarly, my blog is the result of abundance-based design. Nobody is forcing me to write on a regular schedule. And I have no set limit on the time I spend on my blog.

Many people choose to simulate scarcity-based design, maybe because it comes with an adrenaline rush. In fact, the adrenaline rush is good indication that you are in scarcity mode. You will often hear scarcity-based designers say that they are running of time, money or space. They may spend much time planning or worrying about costs and deadlines. There are many examples of artificial scarcity-based design:

  • One of the great fallacies of software engineering is that what matters in the software industry is how long it takes and how much it costs. But anyone who has been in the software industry long enough knows that the real problem is that most software is bad. Some of it is atrocious. For example, Apple iTunes is a disgrace.  I don’t care whether the iTunes team finished on time and within budget. Their software is crap. They failed as far as I am concerned.
  • Nobody cares how long it took  you to write your novel or research paper. Yet people sign deals with publishers with fixed deadlines and others choose to publish in conferences with fixed deadline. They create external pressure, on purpose.

Frankly, if you are a designer such as an artist, a fiction writer, a scientist or a scholar, you should have a feeling of urgency, not worry. A single strategy may suffice to put you in abundance mode:

  • Reduce the quantity. Apple is well known for having few products. Despite having billions of dollars, they focus on few projects. And their new project have often fewer features than the competition. By focusing your attention, you ensure abundance. Don’t start more projects than you can’t execute with ease.

Further reading: Publishing for Impact by John Regehr and The merits of chasing many rabbits at the same time by Alain Désilets.

Is science more art or industry?

picture by bdesham
In my previous post, I argued that people who pursue double-blind peer review have an idealized “LEGO block” view of scientific research. Research papers are “pure” units of knowledge and who wrote them is irrelevant.

Let us take this LEGO block view to its ultimate conclusion.

If science is pure industry, producing standardized elements—called research papers, why should papers be signed as if they were pieces of art? The signature is obviously irrelevant. Nobody cares who made a given LEGO block. Thus, I propose we omit names from research papers. It should not change anything, and it will be fairer.

Indeed, why not have anonymous papers all the way? Journals could publish articles without ever telling us who they are from. We would ignore, for example, which papers were written by Einstein or Turing. How is that relevant? How does it help us to appreciate a given paper to know it was written by Turing?

What would we do for conferences? Because papers are standard units, people could attend conferences and be assigned a paper, any paper, to present. Presenting your own work is a bit too egotistical anyhow.

Of course, for recruiting or promotion purposes, we would need to be able to map research papers to individuals. But, because papers are standard units, all you care about is the number of papers and related statistics. Thus, an academic c.v. would not list research papers, but instead provide a key that could be used to retrieve productivity statistics.

Of course, this is not, even at a first approximation, how science works. Science is more art than industry. That is why we put our names on research papers. It does matter that it is Turing that wrote a given paper. It helps us understand the paper better to know its author, its date and its context. When I receive a paper to review, I try to see how the authors work, what their biases are.

Research papers present a view of the world. But like Plato’s cave, this view is fundamentally incomplete. If a paper report the results from some experiments they conducted, the paper is not these experiments: it is only a view on these experiments. It is necessarily a biased view. Do you know what the biases of these particular authors are?

Let us be candid here. When reviewing research papers, there is no such thing as objectivity. Some papers are interesting to the reviewer, some aren’t. What makes it interesting has to do with whether the world view presented is compatible with the reviewer’s world view. And because different individuals have (or should have) different world views, it does matter who wrote the paper even if we omit names. It helps me to find your paper interesting if I can put myself in your shoes, get to know who you are. An anonymous paper is far more likely to be boring to me, because it is hard to have empathy for the authors.

Some days, we all wish it did not matter who we are. Can’t people just look at our work on its own? You can get your wish by becoming a bureaucrat or a factory worker. Science is for people who want to see their name in print, people who want to build their reputation and cater to their inflated ego. In short, good science is interesting.

The case against double-blind peer review

Many scientific journals use double-blind peer review. That is, the authors submit their work in a way that cannot be traced back to them. Meanwhile, the authors do not know who the reviewers are. In this way, the reviewers are free to speak their mind. It feels fair because the reviewers cannot be influenced (in theory) by the declared affiliation of the authors or their relative fame.

How well does it work in practice? You would expect double-blind reviewing to favor people from outside academia. Yet Blank (1991) reported that the opposite is true: authors from outside academia have a lower acceptance rate under double-blind peer review. Moreover, Blank indicates that double-blind peer review is overall harsher. This is not a surprise: It is easier to pull the trigger when the enemy wears a mask.

Meanwhile, there is at best a slight increase in the quality of the papers due to double-blind peer review (De Vries et al., 2009), everything else being equal. However, not everything is equal under double-blind peer review. What is the subtext? That somehow, the research paper is a standalone artefact, an anonymous, standardized piece of LEGO. That it should not be viewed as part of a stream of papers produced by an author. It sends a signal that an original research program is a bad idea. Researchers should be interchangeable. And to assess them, we might as well count the number of their papers since these papers are standard artifacts anyhow.

But that is counter-productive! Research papers are often only interesting when put in a greater context. It is only when you align a series of papers, often from the same authors, that you start seeing a story develop. Or not. Sometimes you only realize how poor someone’s work is by collecting their papers and noticing that nothing much is happening: just more of the same.

Researchers must make verifiable statements, but they should also try to be original and interesting. They should also be going somewhere. Research papers are not collection of facts, they represent a particular (hopefully correct) point of view. A researcher’s point of view should evolve, and how it does is interesting. Yet it is a lot easier to understand a point of view when you are allowed to know openly who the authors are.

Are there cliques and biases in science? Absolutely. But the best way to limit the biases is transparency, not more secrecy. Let the world know who rejected which paper and for what reasons.

Source: This blog post came about through an online exchange with Philippe Beaudoin.

References:

  • Blank, R.M., The effects of double-blind versus single-blind reviewing: Experimental evidence from the American Economic Review, The American Economic Review 81 (5), 1991.
  • De Vries, D.R. and Marschall, E.A. and Stein, R.A., Exploring the Peer Review Process: What is it, Does it Work, and Can it Be Improved? Fisheries 34 (6), 2009.

Update: Mark Wilson has another argument against double-blind peer review. What if you pick up good ideas from double-blind papers that are later rejected and remain unpublished? How do you acknowledge the contribution of the authors of the unpublished work?

Update 2: Patrick Lam points out that the programming languages community now uses a variant of double-blind review for some conferences (like PLDI or POPL, the top PL conferences) where the authors are asked to submit blinded papers, but the identities are revealed to the reviewers after they submit their first-draft reviews.

Further reading: I have more comprehensive argument in a latter blog post.

Ten things Computer Science tells us about bureaucrats

Originally, the term computer applied to human beings. These days, it is increasingly difficult to distinguish reliably machines from human beings: we require ever more challenging CAPTCHAs.

Machines are getting so good that I now prefer dealing with computers than bureaucrats. I much prefer to pay my taxes electronically, for example. Bureaucrats are rarely updated, and they tend to require constant attention like aging servers.

In any case, a bureaucracy is certainly an information processing “machine”. If each bureaucrat is a computer, then the bureaucracy is a computer network. What does Computer Science tell us about bureaucrats?

  1. Bureaucracies are subject to the halting problem. That is, when facing a new problem, it is impossible to know whether the bureaucracy will ever find a solution. Have you ever wondered when the meeting would end? It may never end.
  2. Brewer’s theorem tell us that you cannot have consistency, availability and partition tolerance in a bureaucracy. For example, accounting departments freeze everything once a year. This unavailability is required to achieve yearly consistency.
  3. Parallel computing is hard. You may think that splitting the work between ten bureaucrats would make it go ten times faster, but you are lucky if it goes faster at all.
  4. One the cheapest way to improve the speed of a bureaucracy is caching. Keep track of what worked in the past. Keep your old forms and modify them instead of starting from scratch.
  5. Pipelining is another great trick to improve performance. Instead of having bureaucrats finish the entire processing before they pass on the result, have them pass on their completed work as they finish it. If you have a long chain of bureaucrats, you can drastically speed up the processing.
  6. Code refactoring often fails to improve efficiency. Correspondingly, shuffling a bureaucracy is just for show: it often fails to improve productivity.
  7. Bureaucratic processes spend 80% of their time with 20% of the bureaucrats. Optimize them out.
  8. Know your data structures: a good organigram should be a balanced tree.
  9. When an exception occurs, it goes back the ranks until a manager can handle it. If the CEO cannot handle it, then the whole organization will crash.
  10. The computational complexity is often determined by looking at the loops. That is where your code will spend most of its time. In a bureaucracy, most of the work is repetitive.

Update: Neal Lathia commented that neither bureaucrats nor computers understand humor.

Update: “This is a fairly well-known model, and no it isn’t computer science that is at the root of what you are noticing. It is early operations research. Taylorism in fact. There was a conscious effort in the 20s and 30s to bring Taylorist style a…ssembly line/operations research thinking into white collar work, starting with organizing pools of typists, secretaries and other office workers the same way banks of machine tools were organized into flow shops and assembly lines. The exact same Taylorist time-and-motion study tools were applied (in fact, in the 30s this was so popular that women’s magazines carried articles about time-and-motion in the kitchen. Example: puzzles like “what’s the fastest way to toast 3 slices of bread on a pan that can hold 2 and toast 1 side at a time?) Computer science itself was initially strongly influenced by shopfloor OR… that’s where metaphors like queues come from after all.” (Venkatesh Rao)

The Open Java API for OLAP is growing up!

olap4j log
Software is typically built using two types of programming languages. On the one hand, we have query languages (e.g., XQuery, SQL or MDX). On the other, we have the regular programming languages (C/C++, Java, Python, Ruby). A lot of effort is spent on the mismatch between these two programming styles. It remains a sore point in many projects.

Microsoft has been trying especially hard to resolve this mismatch. Their LINQ component allows you to use relational or XML data sources directly in your favorite language (e.g., C#).

Oracle has its own solution, the Oracle Java API. It allows you to query OLAP databases directly from Java, without SQL or MDX.

Unfortunately, these solutions are vendor-specific. With the rise of Open Source Business Intelligence, we seek open solutions which are shared and co-developed.

That is what the Open Java API for OLAP (olap4j) is. Anyone can build an OLAP engine and offer support for olap4j. You then get MDX support for free. Best of all, an application written against olap4j should work with any olap4j-compliant OLAP server which includes SQL Server Analysis Service and SAP Business Information Warehouse. And if you add Mondrian, you can get olap4j compliance out of any common relational database management system such as MySQL.

Lead by the Linus Torvalds of OLAP (Julian Hyde), olap4j finally reached version 1.0. A leading-edge feature I find interesting is that olap4j supports notifications which should enable real-time OLAP applications. Whenever something changes at the database level, the OLAP server can notify its clients, effectively pushing a notification. One obvious application is in the financial industry where data must be quickly updated.

Further reading: The press release for olap4j 1.0. Julian Hyde’s blog post on this topic. Luc Boudreau’s blog post. See also some of my older blog posts: JOLAP is dead, OLAP4J lives? (2008) and JOLAP versus the Oracle Java API (2006).

How information technology is really built

One of my favorite stories is how Greg Linden invented the famous Amazon recommender system, after after being forbidden to do so. The story is fantastic because what Greg did is contrary to everything textbooks say about good design. You just do not bypass the chain of command! How can you meet your budget and deadline?

In college, we often tell students a story about how software and systems are built. We gather requirements, we design the system, we get a budget, and then we run the project, eventually finishing within budget and while respecting the agreed upon time frame.

This tale makes a lot of sense to people who build bridges, apparently. It not like they can afford to build three different bridge prototypes and then ask people to choose which one they prefer, after checking that all of them are structurally sound.

But software systems are different.

Consider Facebook. Everyone knows Facebook. It is a robust system. It serves 600 million users with only 2000 employees. Surely, they are excessively careful. Maybe they are, but they do not build Facebook the way we might build bridges.

Facebook relies on distributed MySQL. But don’t expect any 3 Normal Forms. No join anywhere in sight (Agarwal, 2008). No schema either: MySQL is used as a key-value store, in what is a total perversion of a relational database. Oh! And engineers are given direct access to the data: no DBA to preserve the data from the evil and careless developers.

Because they don’t appear to like formal conceptual methodologies, I expect you won’t find any entity-relationship (ER) diagram at Facebook. But then, maybe you will find them in large Fortune 100 companies? After all, that is what people like myself have been teaching for years! Yet no ER diagram was found in ten Fortune 100 companies (Brodie & Liu, 2010). And it is not because large companies have simple problems. The average Fortune 100 has ten thousand information systems, of which 90% are relational. A typical relational database has between 100 and 200 tables with dozens of attributes per table.

In a very real way, we have entered a post-methodological era as far as the design of information systems is concerned (Avison and G. Fitzgerald, 2003). The emergence of the web has coincided with the death of the dominant methods based on the analytic thought and lead to the emergence of sensemaking as a primary paradigm.

This is no mere coincidence. At least, two factors have precipitated the fall of the methodologies designed in the seventies:

  • The rise of the sophisticated user. These days, the average user of an information system knows just as much about how to use the systems than the employees of the information technology department. The gap between the experts and the users has fallen. Oh! The gap is only apparent: few users even understand how the web work. But they know (or think they do) what it can do and how it can work. Yet, we continue to see users as mere faceless objects for who the systems are designed (Iivari, 2010). The result? 93% of accounts are never used in enterprise business intelligence systems (Meredith and O’Donnell, 2010). Users now expect to participate in the design of their tools. For example, Twitter is famous for its hashtags which are used to mine trends, and which are the primary source of semantic metadata on Twitter. Yet did you know that they were invented by a random user, Chris Messina, in a modest tweet back in 2007? It is only after users started adopting hashtags that Twitter, the company, adopted it. Hence, Twitter is really a system which is co-designed by the users and the developers. If your design methodology cannot take this into account, it might be obsolete. Recognizing this, Facebook is not content to test new software in the abstract, using unit tests. In fact, code is tested during the deployment for user reactions. If people react badly to an upgrade, the upgrade is pulled back. In some real way, engineers must please users, not merely satisfy formal requirements representing what someone thought the users might want.
  • The exploding number of computers. According to Garner, Google had 1 million servers in 2007. Using cloud computing, any company (or any individual) can run software on thousands of servers worldwide without breaking the bank. Yet Brewer’s theorem says that, in practice, you cannot have both consistency and availability (Gilbert and Lynch, 2002). Can your design methodology deal with inconsistent data? Yet, that is what many NoSQL database systems (such as Cassandra or MongoDB) offer. Maybe you think that you will just stick with strong consistency. JPMorgan tried it and they ended up freezing $132 million and losing thousands of loan applications during a service outage (Monash, 2010). Most likely, you cannot afford to have strong consistency throughout without sacrificing availability. As they say, it is mathematically impossible. Brewer’s theorem is only the tip of the iceberg though: what works for one mainframe, does not work for thousands of computers. Not anymore than a human being is a mere collection of thousands of cells. There is a qualitative difference in how systems with thousands (or millions) of computers must be designed compared with a mainframe system. Problems like data integration are just not on your radar when you have a single database. We have moved from unicellular computers to information ecosystems. If your design methodology was conceived for mainframe computers, it is probably obsolete in 2011.

Building great systems is more art than science right now. The painter must create to understand: the true experts build systems, not diagrams. You learn all the time or you die trying. You innovate without permission or you become obsolete.

Credit: The mistakes and problems are mine, but I stole many good ideas from Antonio Badia.

References:

You can assess trends by the status of the participants

I conjecture that, everything else being equal, the level of your education is inversely correlated with innovation.

  • At first, a new idea appears interesting, but it carries no prestige. And there are few financial incentives. Think homebrew computers before Apple. Or blogging in 2003. The people who first join are sociopaths (as per the Gervais principle) who often lack formal education. They recognize what this new idea might do. Yet they are unconcerned by their place in society. They may not even have a resume.
  • Once an idea picks up steam, incentives become more apparent. The community then expands to include people who are slightly more conformist. You will start to see more college degrees. Companies are built. Jobs are created. Think blogging in 2005, or Apple releasing its first personal computer.
  • After some time, it has become obvious that the idea is solid. Think Apple a few months after the Apple II was launched. Or blogging in 2010. People who value greatly prestige finally join up. You start to see prestigious degrees. There are now established practices and some level of expected conformity.

If my conjecture is at least partly true, then you can assess trends by the status of the participants. When you see new trends, such as homebrew 3D printers or open source electronics, how many MIT degrees do you see? Conversely, by the time the prestigious degrees are flocking in, maybe the real innovation is elsewhere?

Fun fact: In 2007, Morgan Stanley, Lehman Brothers, JP Morgan, and Goldman Sachs were among the top 10 employers of MIT graduates. (Source)

Acknowledgement: I was inspired to write this post by P. Bannister.

Social Web or Tempo Web?

Back in 2004, Tim O’Reilly observed that the Web had changed, and coined the term Web 2.0. This new Web is made of several layers which enable the Social Web. Wikipedia and Facebook are defining examples of the Social Web.

This sudden discovery of the Social Web feels wrong to me. In the early nineties, I was an active user of Bulletin Board Systems (BBS). While it was not the Web, or even part of the Internet, BBSes were clearly a social media. You know the multi-user games people play on Facebook? We had that back in 1990. The graphics were poorer, obviously, but it was all about meeting people.

The barrier to entry keeps getting lower, to the point were even grand-fathers are now on Facebook. But the Web has hardly been limited to an elite. Even BBSes were quite democratic: retired teachers would chat with young hackers all the time. It is the extreme low cost of computers and their ubiquity which makes the Social Web so widespread.

A much more interesting change has received less notice: the tempo of the Web is changing. Geocities made it easy for anyone to create a home page. But updating your home page was a slow process. In effect, our mental model of the Web was that of a library, and Web sites were books that could be updated from time to time. Eventually, we gave up on this model and decided to view the Web as a data stream. This realization changed everything.

The pace used to range from static web pages to flaming on posting boards. We have now expanded our temporal range. We can now communicate with high frequency in short bursts. Twitter is one extreme: it is akin to techno music. Facebook is somewhat slower, and more elaborate, maybe  like rock. Posting research articles is no music at all: it is akin to the rhythm of the Earth around the Sun. These tools don’t just differ on the frequency of the updates, but also on their volume, and on the length of the pauses.

Maybe we should try to understand the Web by analogy with music. How does the Web sound to you, today?

Further reading: See my blog post Is Map-Reduce obsolete? Also, be sure to check Venkatesh Rao’s blog. Rao has a new book which I will review in the future.