← All writing

8 min read Draft

Three architectural choices that mattered

What I learned wiring custom telemetry into a multi-machine production line — three calls that paid off, and one that didn't.

A short case study showing what the writing pipeline looks like end to end. Real posts will be 1,000–1,500 words and built around a single production decision and its follow-on effects.

1. Subscribe, don’t poll

Polling OPC-UA is a footgun. Notifications scale with rate-of-change; polling scales with naivety.

async with Client(url=opcua_url) as client:
    nodes = [client.get_node(nid) for nid in node_ids]
    handler = SubHandler(emit=publish_to_bus)
    sub = await client.create_subscription(period=200, handler=handler)
    await sub.subscribe_data_change(nodes)
    await asyncio.Event().wait()

A 200ms publishing interval with deadband on the server side gave us roughly an order of magnitude reduction in bus traffic versus the polling implementation it replaced — at zero cost to UI latency.

2. Buffer at the gateway, not the UI

The first thing operators do on a fresh dashboard is open it on a flaky tablet. If the UI is the durability boundary, you’ve already lost.

Treat the dashboard as a read replica, not a journal.

Postgres + TimescaleDB on the gateway gave us a fast, queryable replica of recent state. WebSockets streamed the live edge; HTTP queries fetched the rest.

3. The wrong call: a single binary for the gateway and the API

We initially shipped one process for both the OPC-UA gateway and the read API. It made operations easier — one container, one log stream — until the day we needed to redeploy the API and a transient hiccup dropped the OPC-UA subscriptions for six minutes.

The fix was boring and obvious in hindsight: split them. They share a database; they share nothing else. Independent restarts cut downtime to zero.