On 1 October 2026, between 13:27 and 13:31 IST (07:57 to 08:01 UTC), the Testsigma web application in the US region was unavailable for approximately four minutes. A burst of requests from a single API integration exhausted the database connection capacity on our application servers, and the automatic health checks then restarted all servers at the same time. The database itself remained healthy, and no customer data was lost, altered or exposed. Service was fully restored by 13:31 IST.
All accounts hosted in the US region were affected for approximately four minutes. The EU and India regions were not affected.
During the window, users of the web application and of the public API may have seen:
Test runs already executing on agents continued; results submitted during the window were accepted once the servers were back.
All times are on 1 October 2026 in IST (UTC+5:30).
The outage was triggered by a sudden, sustained burst of bulk data-export calls from a single customer integration, which saturated the application tier's database connection capacity within seconds. Four contributing factors, each benign in isolation, compounded into a brief region-wide interruption.
The database tier itself remained healthy throughout: load, locking and response times stayed within normal ranges, no failover occurred, and no customer data was lost, altered or exposed. No hardware fault or security event was involved.
Service recovered within four minutes once the servers restarted. Full recovery was confirmed at 13:31 IST and the platform has been monitored for stability since. The investigation across application logs, database metrics and API traffic is complete, the root cause has been identified, and the integration that generated the burst has been identified.
Short-term measures. We have started a set of configuration and monitoring changes that require no code release and will be completed over the coming days. These changes allow the platform to tolerate a temporary database slowdown without restarting application servers, alert our on-call team as soon as database connection usage runs high, and address the specific integration traffic pattern that triggered this incident.
Permanent measures. In parallel, our product and engineering teams are implementing changying causes: per-account limits on API request rates so that no single integration can affect capacity shared by other customers, improvements to the affected API so that reading data does not create additional database work, bounded page sizes on list APIs, dedicated connection capacity for interactive use, and regular load testing against traffic surges. Most accounts will see no difference from the API request limits, which will be set well above normal usage; the limits will be published, and accounts with high-volume integrations notified, before they take effect.