A month ago, Outbid’s viral launch broke Vemetric’s ingestion. Since then, I’ve spent a good amount of time rebuilding the part of Vemetric that turns incoming events into sessions, devices and the data you see in your dashboard.
The short version: Vemetric’s ingestion can now scale horizontally. If there’s more traffic, I can add more workers, instead of watching the queues grow.
In this post I want to walk you through the four main changes, and what I want to work on next.
1. No More Global Concurrency Limit
This was the biggest bottleneck during the incident. Session and device jobs were processed one at a time, no matter how many workers were running.
The limit existed because two workers updating the same session at the same time would overwrite each other’s changes. But it also meant that adding more workers didn’t help at all.
With the changes below, sessions don’t depend on this limit anymore, so I could remove it. Now every worker picks up session and device jobs, and adding a worker actually adds throughput.
2. Buffering Sessions in Redis
A session changes with every page view. It gets longer, and sometimes it gets a location. Before, every one of these changes meant reading the session from ClickHouse and inserting it again.
Now the current state of every active session is buffered in Redis. Workers update it with atomic operations, so parallel updates can’t overwrite each other. Once per second, a single flusher takes all sessions that changed and writes them to ClickHouse in one batch. Every update also increases the session’s revision number, and the new session table simply keeps the highest revision.
3. Devices Without Duplicates
Devices follow the same idea, just simpler. The new device table removes duplicate inserts by itself, so it doesn’t matter if two workers create the same device at the same time.
Also, instead of creating a device job for every single event, the hub now creates at most one per visitor and device every two minutes.
4. Less Work per Event
Besides adding more workers, I also wanted to make each job cheaper. Previously, a single event caused about six database round trips across its event, session and device jobs. Now it’s roughly one batched write:
- Users are cached in Redis for a minute
- The hub passes the project’s domain along, so there’s no need to look up the project
- Session updates happen in Redis
- All writes to ClickHouse are batched, either by the flusher or by ClickHouse’s async inserts
New Tools for Testing the Ingestion
Changes like these are hard to test by just clicking through the dashboard. So as part of this work, I extended Vemetric with a few tools that are now part of the repository and that I can use for every future change to the ingestion:
- Load test: sends traffic to a local hub at a fixed rate and checks that every event was stored.
- Comparison: runs the same scripted traffic through two versions of the code and compares all sessions, devices, users, events and dashboard queries.
- Query benchmark: generates large amounts of test data and measures how fast the dashboard queries are.
- Migration rehearsal: runs a migration against a copy of production, restored from the daily backup.
For this rebuild, the results looked good. At 500 events per second, the old pipeline was still 15 minutes behind after the traffic stopped, while the new one caught up after 10 seconds with just a single worker.
Future Improvements
With these changes, Vemetric’s ingestion is in a much better place. There are a few things I want to work on next:
Upgrading to Bun 1.4
Vemetric’s hub and workers run on Bun, and the upgrade to Bun 1.4 will give both of them a big performance boost. This means more events per second with the same resources, on top of the improvements described above.
Better User Merges
When an anonymous visitor logs in and gets identified, their visit can still end up split into two sessions. I want to improve how anonymous and identified users are merged, so a visit stays one session and your user journeys are more accurate.
Improved Bot Detection
After these changes, Vemetric can handle the load from bot traffic a lot better. But bots with normal-looking browsers still get through and skew your analytics data, which is why I want to filter out this kind of traffic more reliably.
At the same time, some bots are actually interesting to know about, like crawlers, search engines and AI agents visiting your website. Instead of just throwing that traffic away, I’d like to give you insights into which of them are visiting your website.
Final Words
The Outbid incident really hurt, but it also gave me a very clear list of what had to change. I’m happy with how the rebuild turned out, and I’m much more confident now that Vemetric can handle the next viral launch. 🚀
If you have questions about any of this, or ideas for the next steps, feel free to reach out on X or open an issue on GitHub.