Known-failures baseline
Also called known failures, failure baseline, expected failure, xfail.
A known-failures baseline is the recorded set of tests that already fail, so a later run can be compared with that set.
Example
The recorded set on main already includes some failing tests. You change a glossary page. Those same tests still fail, and no additional test fails. The baseline held. If one more test fails, that failure is outside the baseline. If a listed test starts passing and the list is left unchanged, that is outside the baseline too.
Why it matters
A test suite can already contain failures. The baseline is how you tell a new failure from an old one. pytest, for example, can mark a test you expect to fail, and it counts those expected failures separately from unexpected ones. A baseline can be marks like that, or a written list, or a count of failures you have agreed not to grow. The record changes only when someone updates it on purpose.
A run that matches the baseline is not a claim that every test passed. It is a claim that the run did not add failures, and that the failures already on the list still fail. When one of those starts passing, the list is updated on purpose.
How it shows up on bot.ski
On bot.ski, the unit-test check compares each run with a list of test names kept in the repositoryA repository is a project's stored files and their recorded change history, commonly managed with Git.. The check fails when a test fails that is not on the list, and it fails when a listed test starts passing, so someone updates the list on purpose. Visitors do not see that list on the public pages.
Common confusion
A known failure is not a success. The baseline says the failure was already recorded. It does not mean the test should be ignored forever, and it does not mean a new failure is acceptable.
Sources
Updated October 2, 2026.
