This is a transcript of the video above. You can try WikiWatch live at wikiwatch.io, and the source code is on GitHub.

I was browsing Wikipedia at the weekend, and I stumbled on one of their special pages that, on the surface, looks kind of admin-y. But I saw potential. I ended up building something I think is quite fun, so I want to show you what I built and take you through how it's constructed. Maybe we'll learn something along the way.

Have a look at this. This is Wikipedia's Recent Changes page. It's a list of what's been changed in the past few minutes, and there's a huge number of pages that people have been editing.

Wikipedia's Recent Changes page: a dense list of the latest edits, each with its time, article title, size change and editor

On the surface this is kind of dull, and I think that's okay. It's not supposed to be exciting. It's supposed to be a log file, and it does that job well. But I saw something in this. I thought if I could get hold of this information and manipulate it myself, I could build something a lot more compelling.

So my first draft looks like this. I just grabbed everything that was on the edits list and started playing it back on my own server.

WikiWatch's Every Edit page: the latest 100 Wikipedia edits, newest first, with time, size change, article and editor

From that I was able to start crunching the numbers and come up with a leaderboard. This is the leaderboard of who's editing what, and what's the most popular thing to edit at the moment. It's currently the All India Football Federation.

WikiWatch's Most Active page: the All India Football Federation is the most-edited article, with its logo, summary and a timeline of recent edits, beside a live feed of the latest edits

One thing I've learned looking at this: there's always plenty of sport in the top 100. There's Gérard Depardieu. I'm not going to touch him. I'm not going to go near that article. Self-Portrait with Death Playing the Fiddle. Who knew that was popular?

Further down the leaderboard: Gérard Depardieu, the Duke of Northumberland and Self-Portrait with Death Playing the Fiddle, each with a preview image and edit timeline

One thing I find fascinating about this is that really old questions, which you'd think would be long settled, are currently being edited. September the 8th. I wonder what happened. We'll find out. I'll dig into that later.

I published it on the internet and sent it to a friend of mine. He looked around, and eventually he saw an article where people were arguing over it. Because you get this thing on Wikipedia: someone edits an article, someone disagrees with them and rolls their changes back. So they make the changes again, and someone rolls them back again, and you get this little bun fight of people arguing over a Wikipedia page.

Turns out you can detect this. And so my third and final page is the list of active arguments on Wikipedia.

WikiWatch's Active Arguments page: two editors repeatedly reverting each other on the Alcohol intoxication article, drawn as a zig-zag of coloured dots, one per revert

People are currently arguing about alcohol intoxication. Not too surprising. De Linke, whatever that is. There's an argument about Donut Lab. The Redeemed Zoomer. And Kansas seems to be controversial this morning. I wonder why.

More active arguments: Bushido (rapper), and a long back-and-forth over Kansas with 21 reverts

I think it's great fun. It's fun to look through and pick up bits of trivia you didn't know you were interested in.

I want to take some time and show you how this is built with Spacetime, and the general architecture, so if you want to build something similar yourself, you can. If you just want to see it, head to wikiwatch.io. It's live now. But for now, let's get stuck into the architecture.

Fundamentally, the architecture for this isn't complicated. We just go to Wikipedia periodically, grab the latest data and store it somewhere, and then we want to ship the live changes over to the front end. That corresponds very neatly to the three big ideas in Spacetime. Three big ideas, three simple ideas: functions, tables and subscriptions.

Animated architecture diagram: Wikipedia's recent changes feed a periodic fetch function, which stores edits in an edit table, which ships live changes to a React frontend; the three middle stages are labelled as Spacetime's three ideas: function, table and subscription

So let's break down the first Wikipedia special page, and you'll see how the end-to-end pipeline goes.

Going back to the Wikipedia page, they've got this list of edits. Luckily they also publish it as JSON, a structured data version, which looks like this.

So all I did was find the edit object I'm interested in and convert it into a Spacetime table definition by changing the values into the types of those values. Hit spacetime publish, and I've got a table ready to store the data. That part's pretty mechanical.

Animated diagram, 'From JSON to a table': one edit object from Wikipedia's recent changes JSON is highlighted and copied into a Spacetime table definition, each literal value turns into its type in turn (t.u64, t.string, t.bool, t.timestamp), then spacetime publish creates an empty edit table ready to store the data

Next I need a function that's going to periodically go to Wikipedia and get the data. This is fairly straightforward stuff. It does an HTTP fetch and parses the JSON out, and now I've got a list of edits to put into my table. Open a transaction to the table and insert each one.

Animated diagram of the fetchRecentEdits procedure: as each line of code appears, it fetches Wikipedia's recent changes, parses the JSON into a batch of edits and inserts each one into the edit table; a second, overlapping batch arrives, the edits already stored are caught by a duplicate check and skipped, and only the new ones are inserted

The only thing that makes this slightly complicated is that I'm checking for new edits on Wikipedia periodically, so there are going to be overlaps. There are going to be edits in there that I already saw last time. So I do a check before I insert to make sure it's not already in the database, and if it is, I skip it.

Now if I run that function regularly, I've captured all the information that Wikipedia has, and I can start displaying it.

To do that, I need a half extra concept: a scheduler. They're really easy to set up in Spacetime. You need to decide how often it's running and store that somewhere. Where do you store that information? In a table. All data is stored in a table in Spacetime.

So you create a table that's going to store how often we run the fetch process, and then you point the function at that table, like this. And that does it. That will now run every 15 seconds.

Animated diagram of the scheduler: the interval 'every 15 seconds' is stored as a row in a schedule_fetch_recent_edits table, defined with a scheduled_at column; the fetchRecentEdits procedure points at that table with onSchedule, and a timeline shows it firing every 15 seconds

So once I've got that going, I'm getting a feed of edits through to the backend. How do I display that on the front end?

This is one of my favourite parts of Spacetime, because it's super easy. You want something that will run in the front end (I'm using React for this) that will go off to the database, run a query, get the latest stuff, and then periodically be updated with the new stuff.

Actually, what you really want is to run the query, get the stuff, and then just be told when there's anything extra to add.

Actually, what you really want is to just say what you're interested in, and have it magically appear on the front end. And that's what Spacetime does.

Animated comparison of ways to get database edits to a React frontend: polling with repeated queries is crossed out, querying then being pushed new rows is crossed out, and simply declaring what you want wins; then a single line, useTable(tables.edit), mirrors the server's edit rows into the browser over a WebSocket, with new edits flying across as they arrive

In React land, you just call useTable, supply your query, and the front end will magically have whatever the backend has. As new updates come in, they'll be shipped over a WebSocket, but you don't even need to worry about that. It's just live updated, and you can get on with the problem of rendering.

The code for this is super simple. This is what initially attracted me to Spacetime. If you want to keep the client in sync with what the server sees, it's trivial. The work is done for you.

So that's enough to give me my edit page and start live streaming data. And if we step back and take a look at that, that's pretty much the architecture that runs through this whole thing.

Animated recap of the WikiWatch architecture: a 15-second schedule triggers the fetchRecentEdits function, which writes Wikipedia edits into the edit table, and a useTable subscription streams them to the React frontend, with packets flowing end to end along the pipeline

But it gets a bit more complicated. Not architecturally more complicated, but operationally more complicated. Let me show you what I mean.

I had this up and running for a day or so, and I started to realise quite how many edits Wikipedia gets in a day. It's something like 100,000 rows a day. Now, I could store that many rows, and that would be okay for the database. I could ship them to the front end, but it gets sluggish, and that worries me more. I don't really want to do any of that. What I really want is to keep a live set of edits to look at and manipulate.

So I decided to do some filtering and pruning, with two extra jobs. One of them runs every few hours and says, "You know what? We don't need any edits that are older than a day, so delete them." That keeps the working set to about 100,000 rows.

Animated diagram, Filtering and pruning: two scheduled jobs appear, deleteOldHistory and expireOldEdits; focusing on deleteOldHistory, a timeline of edits marks everything older than one day, those edits are deleted, and the working set settles at about 100,000 rows

The other one is slightly more complicated. Even though I'd quite like to keep a day's worth of data, I don't really want to ship a day's worth of data to the front end, because it's not going to display that much. So I made a slight enhancement to my edit table. I added a live flag that says whether this is something that should be shown on the front end. Every new edit comes in as live, and then I set up another simple schedule that says, "If it's older than 15 to 30 minutes, flip that flag to false."

Then on the front end, my useTable goes from "grab all the edits" to "grab all the live edits", and that's all there is to it. That trims the front-end set down to, I think, about 1,000 to 1,500 rows usually, which is much more manageable and much snappier on that first page load.

Animated diagram, A live flag: shipping a whole day of edits (about 100,000 rows) from Spacetime to the browser is blocked; instead the edit table gains a live column, new edits arrive live, a scheduled job expireOldEdits sets live to false on edits older than about 15 minutes, and the client's useTable(tables.edit) gains .where(e => e.live.eq(true)); the result is about 100,000 rows stored but only 1,000 to 1,500 shipped, for a much snappier first page load

The last part of this is that edits on their own look a bit dull. What you want is the flavour. You want the image, and some text telling you what you're looking at. So I also had to go and fetch previews for all these articles.

That again is the same pattern. I found a JSON endpoint, turned it into a table in Spacetime, and then wrote a fetcher.

Animated diagram, 'The same pattern': chips for JSON endpoint, table and fetcher light up in turn as the Wikipedia API's JSON preview for an article (page id, title, extract, thumbnail, description) appears beside the article_preview table definition built from it

The only thing that makes this fetcher slightly more interesting is deciding which articles to fetch. That's a query. So I open a transaction to the database and ask which edits are live but don't already have a preview. Then I do the fetch, and then another transaction that inserts the previews I've grabbed.

Lastly, I go to the front end, and instead of using a table directly, I use useTable with a join to figure out which previews I want to see for the edits that are relevant.

Animated diagram of the preview fetcher: a scheduled procedure's code lights up in three steps (a transaction finds live edits with no preview, an HTTP fetch asks Wikipedia for previews, a second transaction inserts them), then the frontend's useTable subscription joins live edits to article_preview on page_id so only previews for live edits are shown

And that's pretty much the whole thing. It's all built from three and a half pieces. You've got functions, which grab the data; tables, which store the data; subscriptions, which look at the data; and schedules, which are really only functions that point at tables telling them when to run.

Animated summary, 'Three and a half pieces': cards appear for Functions (grab the data), Tables (store the data) and Subscriptions (look at the data), then a half-width Schedules card with a ticking clock, captioned 'Schedules are just functions pointed at a table that says when to run'

If you want to check it out live, it's on wikiwatch.io. If you want the source code, that's on GitHub. And if you've got any questions, or there's anything you'd like more detail about, leave a comment on the video and I'll get back to you.

Thanks for reading!