How it scales
Two numbers matter for an ingest pipeline: what the database can absorb, and what one process can push into it. We've measured both.
| Measurement | Result | Evidence |
|---|---|---|
| Postgres ceiling | 4,159 write transactions/s; 31,481 alert-list reads/s — benchmarked on a dedicated 16-vCPU test host. | pgbench baseline, September 2026 |
| Single-process sustained ingest | 400 alerts/s with readers, writers, SSE, and reping active — 251,996/251,996 accepted, zero dropped iterations or HTTP failures, 123 ms ingest p95. | mixed-load qualification, October 1, 2026 |
| Burst absorb (1,000 alerts/s) | One app process sustained a two-minute 1,000/s burst: 222,021/222,021 creates accepted across the full run, zero dropped iterations or HTTP failures, 83 ms ingest p95. Measured on a dedicated 24-core 2013-era Xeon host with PostgreSQL 17. | mixed-load qualification, October 1, 2026 |
| Vertical resize | Rehearsal in progress. Targets under test: ~22 s hypervisor resize, ~75 s end-to-end including service recovery. | scheduled — results publish when dated |
| Per-customer silos | Dedicated host + dedicated Postgres per customer, sized for your burst — no noisy neighbors by construction. | Silo SKUs, provisioned by one command |
flowchart LR
subgraph GEN["load generators - separate processes, token-bucket paced"]
L1["loadgen 1"]
L2["loadgen 2"]
L3["loadgen N"]
end
subgraph APP["single-process OpsPing silo"]
P["opspingd
API + event bus + workers"]
end
RL["rate limiter
atomic counters in Postgres"]
PG[("PostgreSQL
separately measured ceiling:
4,159 write tx/s
31,481 alert-list reads/s")]
L1 --> P
L2 --> P
L3 --> P
P --> RL
RL --> PG
Fig 1 — The ingest path. Each silo intentionally runs one modular-monolith process because its event bus and workers are in-process. Capacity comes from sizing that process and its PostgreSQL host, not hiding coordination problems behind replicas.
flowchart TB
subgraph ONE["one loadgen process"]
TB["token bucket paced
at the target rate"] --> WG["concurrent workers"]
WG --> ST["local stats:
accepted / dropped / latency"]
end
ST --> MERGE["merge across processes:
summed throughput, pooled latencies"]
MERGE --> PUB["published numbers -
the server is the bottleneck,
never the load generator"]
Fig 2 — Methodology. Load is paced at the target rate rather than slammed; when one generator process saturates, we add generator processes instead of distorting the measurement.
Rate limiting protects the pipeline from abuse, and the ingest cap is deliberately generous — see the API reference for the exact behavior.
flowchart TB
IN["incoming alerts per second"] --> Q{"Within capacity?"}
Q -- "yes" --> ACK["202 accepted
qualification: 222,021 of 222,021
zero dropped"]
Q -- "over the ingest cap" --> REJ["429 - loud, explicit
rejection, retryable"]
PGDOWN["Postgres unreachable"] --> ERR["5xx back to the sender
never a silent 202"]
Fig 3 — Failure posture under load: overload and infra failure degrade to explicit errors a sender can retry. The pipeline never queues silently and never acknowledges a write it failed to commit.
How it survives failure
Backups you've never restored are hopes, not backups. So we drill.
| Scenario | Behavior | Evidence |
|---|---|---|
| Postgres killed mid-load | Loud failure — no silent data acceptance (every failed write returns an error to the sender). Crash recovery replays the write-ahead log; zero acknowledged-write data loss. Kill-rehearsal results publish when dated. | rehearsal scheduled |
| Provider outage mid-notification | Failed channel sends retry with backoff (1 / 5 / 15 min) from a durable queue; anything still undeliverable dead-letters with an audit entry — it never just vanishes. P1 pushes additionally re-fire until acknowledged. | delivery worker, shipped September 13, 2026 |
| Restore from backup | Prod database dumped, shipped off-box, and restored into a throwaway PostgreSQL 16: 1.26 s restore, verified row-for-row — every table identical to the source. | restore drill, September 13, 2026 |
| Backups | Nightly encrypted dump on-host + an off-box encrypted copy. | documented procedure |
| RPO / RTO | RPO ≤ 24 h today (≤ 5 min once managed Postgres lands — evaluation open); RTO target 45–75 min. | disaster-recovery plan |
The public status page is generated live from the API; an unreachable status page is treated as a signal in itself.
How your alerts stay private
- Tenant isolation is enforced, not promised — every entity carries a tenant id, checks run before team checks, and there is no admin bypass: even OpsPing operators are isolated from customer content by the same middleware.
- Zero standing access — operators reach prod through break-glass procedures, not shared credentials.
- TLS everywhere; backups are encrypted at rest, off-box.
- DNS-only Cloudflare — alert ingestion goes straight to our API; it never transits a third-party edge.
The status page can't be silenced by an outage — it runs on infrastructure fully independent of the app, probes the API from outside our stack, and carries operator updates even if everything else is unreachable.
Full control list: security overview and the security questionnaire (also available on request as a document). Data-protection terms: DPA · subprocessor list.
Your data is yours
Full JSON export of your data — profile, alerts, notification rules, channels — is built in: one authenticated call (GET /v2/export), no support ticket. If you leave, your data leaves with you. API reference.
What we won't claim yet
This section is mandatory and stays current. If it ever disappears, treat the rest of this page as marketing.
- SOC 2 — not pursued; our practices are mapped to the Trust Services Criteria on the SOC 2 page.
- Third-party penetration test — planned; an internal security review is complete and its findings fixed (tracked publicly in our backlog).
- SMS/voice channels — Twilio integration is built; production rollout is in progress.
- Single region today — us-east-2, with cross-region DR planned as part of the managed-Postgres decision.
- Current data-loss window — RPO is bounded by nightly backups (≤ 24 h) until streaming replication or managed Postgres lands.
Commercial
Invoicing with wire/ACH, purchase-order support, and NET 30 on annual contracts. No credit-card-only trap. Request a quote — include your ingest volume and team count and you'll get a concrete number, not a sales call.
A formal SLA document is in progress; until it's signed, the honest commitment is the one this page makes — measured performance, drilled recovery, and a founder's phone number that answers.
Who we are
OpsPing is built by LTFI Tech, LLC (Massachusetts). People ask what LTFI stands for: Learn. Try. Fix. Improve. — the discipline behind everything we ship. About · Contact.
Every figure on this page carries its measurement date. Raw drill logs available on request to prospective enterprise customers — ask via contact. Page last reviewed: 2026-09-13.