Why Web 3.0 will fail

Just when we were beginning to get comfortable with Web 2.0, with its gradients, rounded corners and breathless promises of a better tomorrow, the hype machine has coughed up another shiny new label: Web 3.0. Apparently, this is the future.

The idea is that the dumb old web we know and tolerate will gradually be replaced by something smarter: documents stuffed with semantic information that machines can understand, process and combine into useful new things. Your web pages will no longer merely say something. They will know what they are saying. Or at least that’s the plan.

The great minds of the web have been talking about the Semantic Web for a long time, an eternity if measured in internet years. And, to be fair, the idea is genuinely interesting. Add meaning to all that information floating around out there and suddenly machines can do considerably more with it.

Microformats are already taking us a few tentative steps in that direction, providing simple, open ways of embedding semantic information directly into HTML.

Wonderful. There is, however, one small problem. Humans.

Much of the Semantic Web depends on data being structured correctly. Unfortunately, we have spent the last decade demonstrating that we are spectacularly bad at writing correct HTML. I recently came across a study that examined close to 700 000 web pages. More than 93 percent contained HTML syntax errors. And that’s before we start looking under the carpet for invalid CSS, dubious semantics and whatever horrors somebody pasted in from Microsoft Word.

The dream of enforcing clean, well-formed markup through XHTML has more or less collapsed. Most people writing HTML simply don’t care whether their markup validates. Why should they? Browsers have become extraordinarily talented at swallowing malformed code, figuring out what the author probably meant and displaying something vaguely resembling a web page anyway.

We’ve essentially trained an entire generation of developers to discover that the rules are optional. And now we’re going to build the intelligent, machine-readable Semantic Web on top of that. What could possibly go wrong?

The Semantic Web may very well arrive eventually. Parts of it certainly will. But before machines can reliably understand our documents, humans first have to become considerably better at writing them.

Judging by the current state of the web, I wouldn’t hold my breath.

1 comment

  • avatar
    pogo
    17 Oct, 2009
    hi

Leave a reply