Gremlins mutation testing tracker
Auto-updated weekly by cplieger/ci/.github/workflows/weekly-gremlins.yaml.
Last update: 2026-09-21 08:55
This week: 95.9% efficacy (±0.0% across 2 runs), 98.8% mutant coverage. 10 confirmed live mutants.
Trend: ↗ +3.4% from 12-week mean (92.5%).
Rolling 12-week history
| Run (UTC) |
Mean efficacy |
Stddev |
Mutant coverage |
Live mutants |
Δ efficacy |
| 2026-09-21 08:55 |
95.9% |
±0.0% |
98.8% |
10 |
+0.1% |
| 2026-09-14 08:58 |
95.8% |
±0.2% |
98.8% |
11 |
-0.1% |
| 2026-09-07 08:51 |
95.9% |
±0.0% |
98.8% |
10 |
+0.1% |
| 2026-08-31 07:56 |
95.8% |
±0.2% |
98.8% |
11 |
+0.2% |
| 2026-08-24 05:00 |
95.6% |
±0.2% |
98.8% |
11 |
+0.3% |
| 2026-08-21 15:35 |
95.3% |
±0.0% |
98.7% |
11 |
+3.2% |
| 2026-08-17 04:29 |
92.1% |
±0.2% |
98.7% |
19 |
+0.0% |
| 2026-08-10 04:03 |
92.1% |
±0.2% |
98.7% |
19 |
+0.2% |
| 2026-07-27 04:17 |
91.9% |
±0.0% |
98.7% |
18 |
+2.1% |
| 2026-07-20 04:11 |
89.8% |
±0.2% |
98.6% |
22 |
+1.4% |
| 2026-07-13 02:16 |
88.4% |
±0.2% |
97.7% |
25 |
-1.0% |
| 2026-07-07 23:28 |
89.4% |
±0.0% |
97.5% |
21 |
+2.0% |
Current live mutants (10, this week)
Solid gap — LIVED in all 2 runs (known coverage hole) — 10
httpx.go
- L419 — CONDITIONALS_BOUNDARY
- L428 — CONDITIONALS_BOUNDARY
- L446 — CONDITIONALS_BOUNDARY
- L450 — CONDITIONALS_BOUNDARY
- L563 — CONDITIONALS_BOUNDARY
- L575 — CONDITIONALS_BOUNDARY
- L579 — CONDITIONALS_BOUNDARY
roundtripper.go
- L161 — CONDITIONALS_BOUNDARY
- L326 — CONDITIONALS_BOUNDARY
- L326 — CONDITIONALS_BOUNDARY
How to read
- Mean efficacy: % of runnable mutants killed (or timed-out, treated as caught), averaged across the N runs
- Stddev: variance across runs — high stddev (>3%) signals flaky tests
- Mutant coverage: % of mutants reached by the test suite (test depth)
The "Current live mutants" section is bucketed by how many of the N runs the
mutant LIVED in (N = attempts that week, normally 3):
- Rare flake (1/N): a mutant your tests usually KILL but occasionally let
through. Most actionable — almost always means a flaky test that's not
reliable. Open the section to see the file:line; expect to fix the test, not
the production code.
- Weakly flaky (k/N for 1<k<N): tests catch this mutant inconsistently.
Less ideal than rare-flake but still worth investigating.
- Solid gap (N/N): every run lets this mutant through. Either there's no
test exercising the path, or the test asserts the wrong thing. Action: add a
test or strengthen an existing assertion.
The mutation-regression label is added when this week's mean efficacy drops
5% below the rolling 12-week mean.
A ⚠️ line under "This week" means the number describes the measurement, not the
suite: either the attempts disagreed about which mutants survive (one verdict per
disagreement is false), or the score jumped to a flawless 100% from a week with
live mutants. Equivalent mutants are a permanent noise floor, so a reproducible
100% is not something a test suite can reach.
Free-form notes
Add anything below — won't be touched by the auto-updater.
Gremlins mutation testing tracker
Auto-updated weekly by
cplieger/ci/.github/workflows/weekly-gremlins.yaml.Last update: 2026-09-21 08:55
This week: 95.9% efficacy (±0.0% across 2 runs), 98.8% mutant coverage. 10 confirmed live mutants.
Trend: ↗ +3.4% from 12-week mean (92.5%).
Rolling 12-week history
Current live mutants (10, this week)
Solid gap — LIVED in all 2 runs (known coverage hole) — 10
httpx.goroundtripper.goHow to read
The "Current live mutants" section is bucketed by how many of the N runs the
mutant LIVED in (N = attempts that week, normally 3):
through. Most actionable — almost always means a flaky test that's not
reliable. Open the section to see the file:line; expect to fix the test, not
the production code.
Less ideal than rare-flake but still worth investigating.
test exercising the path, or the test asserts the wrong thing. Action: add a
test or strengthen an existing assertion.
The
mutation-regressionlabel is added when this week's mean efficacy dropsA⚠️ line under "This week" means the number describes the measurement, not the
suite: either the attempts disagreed about which mutants survive (one verdict per
disagreement is false), or the score jumped to a flawless 100% from a week with
live mutants. Equivalent mutants are a permanent noise floor, so a reproducible
100% is not something a test suite can reach.
Free-form notes
Add anything below — won't be touched by the auto-updater.