Degraded availability of the Testsigma Web App, US region

Incident Report for Testsigma Status

Postmortem

Summary

On 1 October 2026, between 13:27 and 13:31 IST (07:57 to 08:01 UTC), the Testsigma web application in the US region was unavailable for approximately four minutes. A burst of requests from a single API integration exhausted the database connection capacity on our application servers, and the automatic health checks then restarted all servers at the same time. The database itself remained healthy, and no customer data was lost, altered or exposed. Service was fully restored by 13:31 IST.

Impact

All accounts hosted in the US region were affected for approximately four minutes. The EU and India regions were not affected.

During the window, users of the web application and of the public API may have seen:

  • Pages loading slowly, timing out, or returning an error.
  • Actions such as saving, starting a run, or fetching results failing at the moment of the restart. These actions were not partially applied; they either completed or did not, and can be retried.
  • API requests returning errors or timing out. Integrations that retry with backoff recovered automatically once service was restored.

Test runs already executing on agents continued; results submitted during the window were accepted once the servers were back.

Timeline

All times are on 1 October 2026 in IST (UTC+5:30).

  • 13:26:55 — An integration belonging to one account begins a bulk data export through the public API at a very high request rate.
  • 13:27:00 — Every application server has run out of database connections. New requests wait for a connection instead of being served.
  • 13:27:54 — The automatic health check, which also needs a database connection, starts timing out on every server.
  • 13:28:06 — Requests waiting for a database connection begin failing with errors.
  • 13:28:54 — After repeated failed health checks, the platform restarts all application servers at once.
  • 13:29:00 — The integration retries its export while the servers are restarting.
  • 13:29:00 to 13:30:30 — New application servers start up.
  • 13:31:35 — Service is confirmed fully restored and stable.
  • 13:31 onwards — Engineering investigates application logs, database metrics and API traffic, and identifies the cause within the day.

Root cause

The outage was triggered by a sudden, sustained burst of bulk data-export calls from a single customer integration, which saturated the application tier's database connection capacity within seconds. Four contributing factors, each benign in isolation, compounded into a brief region-wide interruption.

  1. Unthrottled ingress from a single tenant. The integration drove request volume several orders of magnitude above its established baseline. At the time, the public API enforced no per-account rate controls, so the surge competed on equal terms with interactive traffic from every other account.
  2. Write amplification on a read path. The list operation involved carried a latent inefficiency that caused it to perform persistence work for every record it returned. At normal volumes this was imperceptible; under bulk retrieval it extended connection hold times from milliseconds to seconds.
  3. Undifferentiated connection pooling. Application servers draw on a fixed, shared pool of database connections with no partitioning between bulk, background and interactive workloads. Once the burst occupied the pool, requests from all accounts queued behind it and ultimately timed out.
  4. Health-check coupling. Server liveness checks exercised the same connection pool and therefore failed alongside application traffic. The orchestration layer responded as designed by restarting the affected servers, but because every server was affected simultaneously, a degraded-but-serving state became a short, complete interruption.

The database tier itself remained healthy throughout: load, locking and response times stayed within normal ranges, no failover occurred, and no customer data was lost, altered or exposed. No hardware fault or security event was involved.

Resolution

Service recovered within four minutes once the servers restarted. Full recovery was confirmed at 13:31 IST and the platform has been monitored for stability since. The investigation across application logs, database metrics and API traffic is complete, the root cause has been identified, and the integration that generated the burst has been identified.

Short-term measures. We have started a set of configuration and monitoring changes that require no code release and will be completed over the coming days. These changes allow the platform to tolerate a temporary database slowdown without restarting application servers, alert our on-call team as soon as database connection usage runs high, and address the specific integration traffic pattern that triggered this incident.

Permanent measures. In parallel, our product and engineering teams are implementing changying causes: per-account limits on API request rates so that no single integration can affect capacity shared by other customers, improvements to the affected API so that reading data does not create additional database work, bounded page sizes on list APIs, dedicated connection capacity for interactive use, and regular load testing against traffic surges. Most accounts will see no difference from the API request limits, which will be set well above normal usage; the limits will be published, and accounts with high-volume integrations notified, before they take effect.

Posted Oct 03, 2026 - 09:09 UTC

Resolved

On 1 October 2026, between 13:27 and 13:31 IST (07:57 to 08:01 UTC), users of the
Testsigma web application in the US region may have experienced slow page loads,
errors, or failed API requests for approximately four minutes. Test executions
already running on agents continued, and results submitted during the window were
accepted once service was restored. The EU and India regions were not affected.
No customer data was lost, altered or exposed.

Cause: a sudden surge of bulk API requests from a single customer integration
exhausted the database connection capacity on our application servers. Our
automated health checks then restarted the affected servers simultaneously,
which briefly interrupted service for US-region accounts. The database itself
remained healthy throughout.

Resolution: service recovered once the servers restarted at 13:31 IST and has been
stable since. The root cause has been identified. Short-term changes are being
rolled out so that a similar surge cannot cause servers to restart, and permanent
measures, including per-account API request limits, are being implemented to
prevent a recurrence. A detailed post-incident report is available on request.

We are continuing to monitor the platform closely.
Posted Oct 01, 2026 - 08:00 UTC