All articles

Making Vemetric's Ingestion Horizontally Scalable

A follow-up to the Outbid post-mortem: what I changed in Vemetric's ingestion pipeline, so it can scale horizontally and handle traffic spikes without losing data.

5 min read

A month ago, Outbid’s viral launch broke Vemetric’s ingestion. Since then, I’ve spent a good amount of time rebuilding the part of Vemetric that turns incoming events into sessions, devices and the data you see in your dashboard.


The short version: Vemetric’s ingestion can now scale horizontally. If there’s more traffic, I can add more workers, instead of watching the queues grow.

In this post I want to walk you through the four main changes, and what I want to work on next.

1. No More Global Concurrency Limit

This was the biggest bottleneck during the incident. Session and device jobs were processed one at a time, no matter how many workers were running.

The limit existed because two workers updating the same session at the same time would overwrite each other’s changes. But it also meant that adding more workers didn’t help at all.

With the changes below, sessions don’t depend on this limit anymore, so I could remove it. Now every worker picks up session and device jobs, and adding a worker actually adds throughput.

Session and device jobs were processed one at a time. Now every worker replica takes them in parallel.

2. Buffering Sessions in Redis

A session changes with every page view. It gets longer, and sometimes it gets a location. Before, every one of these changes meant reading the session from ClickHouse and inserting it again.

Now the current state of every active session is buffered in Redis. Workers update it with atomic operations, so parallel updates can’t overwrite each other. Once per second, a single flusher takes all sessions that changed and writes them to ClickHouse in one batch. Every update also increases the session’s revision number, and the new session table simply keeps the highest revision.

Workers merge every update into the session's state in Redis. Once per second, one flusher writes the changed sessions to ClickHouse in a single batch.

3. Devices Without Duplicates

Devices follow the same idea, just simpler. The new device table removes duplicate inserts by itself, so it doesn’t matter if two workers create the same device at the same time.

Also, instead of creating a device job for every single event, the hub now creates at most one per visitor and device every two minutes.

4. Less Work per Event

Besides adding more workers, I also wanted to make each job cheaper. Previously, a single event caused about six database round trips across its event, session and device jobs. Now it’s roughly one batched write:

  • Users are cached in Redis for a minute
  • The hub passes the project’s domain along, so there’s no need to look up the project
  • Session updates happen in Redis
  • All writes to ClickHouse are batched, either by the flusher or by ClickHouse’s async inserts
The work behind one event: about six database round trips before, one batched write now.

New Tools for Testing the Ingestion

Changes like these are hard to test by just clicking through the dashboard. So as part of this work, I extended Vemetric with a few tools that are now part of the repository and that I can use for every future change to the ingestion:

  • Load test: sends traffic to a local hub at a fixed rate and checks that every event was stored.
  • Comparison: runs the same scripted traffic through two versions of the code and compares all sessions, devices, users, events and dashboard queries.
  • Query benchmark: generates large amounts of test data and measures how fast the dashboard queries are.
  • Migration rehearsal: runs a migration against a copy of production, restored from the daily backup.

For this rebuild, the results looked good. At 500 events per second, the old pipeline was still 15 minutes behind after the traffic stopped, while the new one caught up after 10 seconds with just a single worker.

Future Improvements

With these changes, Vemetric’s ingestion is in a much better place. There are a few things I want to work on next:

Upgrading to Bun 1.4

Vemetric’s hub and workers run on Bun, and the upgrade to Bun 1.4 will give both of them a big performance boost. This means more events per second with the same resources, on top of the improvements described above.

Better User Merges

When an anonymous visitor logs in and gets identified, their visit can still end up split into two sessions. I want to improve how anonymous and identified users are merged, so a visit stays one session and your user journeys are more accurate.

Improved Bot Detection

After these changes, Vemetric can handle the load from bot traffic a lot better. But bots with normal-looking browsers still get through and skew your analytics data, which is why I want to filter out this kind of traffic more reliably.

At the same time, some bots are actually interesting to know about, like crawlers, search engines and AI agents visiting your website. Instead of just throwing that traffic away, I’d like to give you insights into which of them are visiting your website.

Final Words

The Outbid incident really hurt, but it also gave me a very clear list of what had to change. I’m happy with how the rebuild turned out, and I’m much more confident now that Vemetric can handle the next viral launch. 🚀

If you have questions about any of this, or ideas for the next steps, feel free to reach out on X or open an issue on GitHub.

Related posts

Keep reading

Ready to understand your users?

Integrate and get valuable insights with Vemetric in minutes.

Start tracking
Pricing About Documentation Customers Changelog Blog