You hear about failures from your users first.

Know what is happening and fix it early.

Eight monitoring views show records moving between systems, jobs and workflows running, queues building and errors as they happen. When something breaks, the assistant reads the failed job and its log, names the likely cause in plain language and proposes a fix for someone to apply. Every change is scheduled, versioned and audited.

A failure seen, explained and fixed before a user noticed A monitoring board shows heartbeats, queue, movement and errors. The queue climbs and a source is throttled; a billing push fails on a timeout; the assistant names the cause (a rate limit after a spike) and proposes a smaller batch and a retry; a person applies it; the retry succeeds, the queue drains and the change is recorded. All normal Strain · provisioning slow, billing throttled Error · billing push failed · seen 09:41:05 Recovered · 09:43:12 · change on the record Heartbeats7 of 7 Queue12480 · rising0 Movement12 480 rows / min Errors030 Throughput · rows per minute Log 09:14:02 · Billing → hub · 240 rows · ok 09:38:10 · Provisioning → hub · 1 120 rows · ok 09:40:31 · hub → Billing · throttled to protect the source 09:41:05 · Job 8 812 · Billing push · failed · timeout 09:43:12 · Job 8 812 · retried 3 of 3 · ok · queue drained Analyse with AI Cause · Billing API rate limit after the 09:40 spike. Fix · dispatch batch 500 → 200, then retry the 3 records. Apply Applied · v42 · recorded in Git · audit chain ✓ Seen, explained and fixed before a user noticed.

Everything that moves is visible

Rows moving between systems, jobs and workflows executing, queues building, a source slowing down, a heartbeat missed: all of it shows on the monitoring screens as it happens. The normal cases scroll past as reassurance. The edge cases stand out the moment they begin, such as a rising queue, a drop in throughput or a step taking ten times longer than usual, long before a user notices anything.

  • Eight live views: heartbeats, queues, movement, throughput, errors, jobs, workflows and portals.
  • Every run logged step by step, and every record's journey traceable.
  • Strain shows as strain, such as rising queues, slow sources and throttling, before it becomes an outage.
Live One wall, every moving part Jobs · queues · workers · latency
You know before the phone rings A monitoring wall shows jobs, queues and workers healthy; one queue's backlog rises, turns amber and is flagged while it is still a trend. Ingress jobs 142 today · all green Egress jobs 96 today · all green Workers 4 nodes · heartbeats ✓ API latency 38 ms median Queue depth: billing egress flagged early Backlog climbing for 20 min, still a trend. seen here first, before any support ticket Every moving part in one glance, all green. One queue drifts, flagged while still a trend.
The whole estate at a glance, with the drift that matters flagged while it is still a trend.

From an error to a fix, with the assistant

When something does break, you get more than a stack trace. The assistant reads the failed job, its log and the mapping behind it, names the cause in plain language and proposes the fix, such as a smaller batch, a truncation rule or a retry with back-off, for someone to apply. The retry runs, the queue drains, and the whole episode is on the record.

  • AI analysis of any failed job, one click away, in plain language.
  • Proposed fixes applied by you, retried automatically, and recorded.
  • Alerts and retries that respect the source system rather than hammering it.
Live From alert to fixed AI reads the log · you approve · retry
From alert to fixed on one screen A failed job turns red; the AI reads the log and proposes a specific fix; you approve; the retry runs green and the queue drains. Job 8114, egress failed ✕ String or binary data would be truncated (column 'Notes') AI read the log Source 'Notes' grew to 500 chars; target is 200, so widen and retry. a proposal, waiting for you Your call Approve fix nothing changes without approval Retry column widened · job 8114 re-run, 3 812 rows queue drained ✓ alert → cause → fix → green, without leaving the screen A job fails and the AI reads the log. A fix is proposed and you approve it. Retry runs green and the queue drains.
The AI reads the log and proposes the fix, you approve, and the retry runs green, all on one screen.

Scheduled, versioned and audited by default

Schedules run with time zones, missed-run policy and overlap prevention. Every configuration change is kept in Git history with one-click rollback that never destroys a populated table. The audit trail is a tamper-evident hash chain, and changes reach production through a paired staging-to-production promotion with a snapshot taken first.

  • Cron scheduling with time zones, missed-run policy and overlap prevention.
  • Git-backed configuration history and job rollback that never destroys rows.
  • A tamper-evident audit chain and staging-to-production promotion with a snapshot before every apply.
Live Scheduled, versioned, audited Scheduled · versioned · audited · promoted
Scheduled, versioned and audited by default Four governance panels light in turn: cron schedules with timezones, git-backed change history with rollback, a tamper-evident audit chain, and a checked staging-to-production promotion. Scheduled: cron and time zone nightly egress · 02:00 Africa/Johannesburg · catch-up on board pack · Mondays 06:00 · delivered as PDF missed windows are caught up later Versioned: git-backed history every config change and every job, committed and diffable rollback never destroys rows "what changed?" always has an answer Audited: tamper-evident chain #4412 · sealed #4413 · sealed #4414 · sealed #4415 · sealed chain verified, nothing altered each entry sealed by the one before it Promoted: staging → production staging ✓ production gates passed ✓ credentials + variables switch automatically missing production value = push refused Scheduled, versioned and audited by default. Promotion passes the same checks each time.
Schedules, git-backed rollback, a tamper-evident audit chain and checked promotion, all on by default.

See it on your own systems.

A demo takes about an hour. Bring the failure that last woke someone up and we will show you how STRAX handles it.