Red CI, explained: root-cause triage for failed GitHub Actions runs
A red check tells you that something failed. It does not tell you what, or why, or what to do next, and the answer is usually a few thousand lines into a log. quellbot reads the log so you can read a sentence.
What it reads
When a GitHub Actions run fails on a repository where quellbot is installed, it collects the pieces a person would open in three tabs: the failed job, the failed step, and the tail of that step's log. Then it reads them together with the thing the log alone cannot show, the diff that triggered the run, and any related files it points at: a migration, a config, a fixture. A stack trace is a symptom. The cause is usually in the change.
What it posts
One comment, three parts: what failed, the likely root cause, and a concrete next step. No log dump, no restatement of the error message. Here is the shape of it, on a failing test:
FAIL upload.spec.ts › rejects a duplicate upload
error: column "hash" does not exist
at uploadOnce (src/uploads/store.ts:41)
Root cause: migration
0043_rename_hashrenameduploads.hashtocontent_hash, butuploadOnce()still selects the old column. Update the query insrc/uploads/store.ts:41and re-run the suite.
The error message said a column was missing. The triage says which migration removed it, which function still uses it, and what to change. That is the difference between a red check and a fix.
CI and CD are both covered
Where the comment lands depends on what failed:
- If the run belongs to an open pull request, the triage is posted on the PR, next to the review, so the author sees it where they are already working. This is the CI case.
- If the run is on a merge commit on your default branch with no open PR, the triage is posted on the commit itself and treated as a deploy failure. This is the CD case, and it is the one that usually pages someone.
A failed deploy gets the same reading as a failed test: the job, the step, the log tail, and the commit that caused it.
Advisory only, by construction
quellbot explains the failure and suggests a fix. It never reruns a workflow, never changes a status check, and never edits your CI. Those are not toggles. Anything under .github/workflows/ is a hard-deny path in the policy gate, and the GitHub App is not granted the workflows permission in the first place, so the platform refuses the write even if the gate did not. A tool that can both diagnose your pipeline and modify it is a tool you have to watch; this one you can leave alone.
If the triage points at a real bug, the fix path is the same as for any issue: open one with a trigger slug, or comment /build, and quellbot plans and builds a pull request for you to review. Or mention @quellbot on the thread with a follow-up question and it answers, citing the file and line.
Where it runs, and what it costs
Like every quellbot run, the triage happens in an ephemeral machine in your own Fly.io account on your own Claude or Codex credential, and the machine is destroyed once the comment is posted. The log and the diff are read there and nowhere else. Triage runs are short and count toward the same monthly cap as everything else, so a flaky pipeline cannot quietly run up a bill.
Turning it on
CI triage is a switch on the Triggers page in the console, per account, with a per-repository override. Turn it on for the repositories where a red check costs you the most time. The next time a run fails, the first comment on it is the one that says why.
Free on every public repository. Install the GitHub App and let the next red check explain itself.
Open the console