Three architectural choices that mattered
What I learned wiring custom telemetry into a multi-machine production line — three calls that paid off, and one that didn't.
A short case study showing what the writing pipeline looks like end to end. Real posts will be 1,000–1,500 words and built around a single production decision and its follow-on effects.
1. Subscribe, don’t poll
Polling OPC-UA is a footgun. Notifications scale with rate-of-change; polling scales with naivety.
async with Client(url=opcua_url) as client:
nodes = [client.get_node(nid) for nid in node_ids]
handler = SubHandler(emit=publish_to_bus)
sub = await client.create_subscription(period=200, handler=handler)
await sub.subscribe_data_change(nodes)
await asyncio.Event().wait() A 200ms publishing interval with deadband on the server side gave us roughly an order of magnitude reduction in bus traffic versus the polling implementation it replaced — at zero cost to UI latency.
2. Buffer at the gateway, not the UI
The first thing operators do on a fresh dashboard is open it on a flaky tablet. If the UI is the durability boundary, you’ve already lost.
Treat the dashboard as a read replica, not a journal.
Postgres + TimescaleDB on the gateway gave us a fast, queryable replica of recent state. WebSockets streamed the live edge; HTTP queries fetched the rest.
3. The wrong call: a single binary for the gateway and the API
We initially shipped one process for both the OPC-UA gateway and the read API. It made operations easier — one container, one log stream — until the day we needed to redeploy the API and a transient hiccup dropped the OPC-UA subscriptions for six minutes.
The fix was boring and obvious in hindsight: split them. They share a database; they share nothing else. Independent restarts cut downtime to zero.