Sifflet - Elevated error rates on aws-prod-eu-west-1-c asset discovery and monitoring workflows – Incident details

Elevated error rates on aws-prod-eu-west-1-c asset discovery and monitoring workflows

Resolved
Major outage
Started about 3 hours agoLasted about 2 hours

Affected

Integrations

Partial outage from 9:06 AM to 9:19 AM, Major outage from 9:19 AM to 9:57 AM, Operational from 9:57 AM to 10:51 AM

aws-prod-eu-west-1-c - Integrations

Partial outage from 9:06 AM to 9:19 AM, Major outage from 9:19 AM to 9:57 AM, Operational from 9:57 AM to 10:51 AM

Datasource monitoring

Partial outage from 9:06 AM to 9:19 AM, Major outage from 9:19 AM to 9:57 AM, Operational from 9:57 AM to 10:51 AM

aws-prod-eu-west-1-c - Datasource monitoring

Partial outage from 9:06 AM to 9:19 AM, Major outage from 9:19 AM to 9:57 AM, Operational from 9:57 AM to 10:51 AM

Updates
  • Resolved
    UTC
    Resolved

    This incident has been resolved. Jobs, including monitors, are running normally.

    Most monitors that should have been processed during the incident window have their status set to "technical error".

  • Monitoring
    UTC
    Monitoring

    We identified the root cause and deployed a fix. All jobs, including monitors and asset discovery, are now running normally. The backlog is being processed.

  • Update
    UTC
    Update

    The database powering job executions is experiencing issues. We're still investigating this problem. Monitors and other jobs, such as catalog discovery, are not running.

  • Investigating
    UTC
    Investigating

    Asset discovery (catalog) and monitors are running with higher latencies and error rates than usual in aws-prod-eu-west-1-c. We're investigating this incident.