A confusion matrix counts actual and predicted labels, exposing different failure types that a single accuracy number can hide.
Python classification metrics: fix label order before counting errors
Operation contract
The review fixture declares label order as zero then one. Rows represent actual labels and columns represent predictions. A false positive marks a clean receipt for review; a false negative misses a receipt that required review. Precision and recall use different denominators, so they describe different costs even when their numeric values happen to match.
Failure and ownership boundary
This sample has five fixture records and no claim about real review quality. Empty positive classes require a declared zero-division policy. Threshold selection, class imbalance, grouped records and uncertainty need a larger evaluation design. Scikit-learn pipelines: fit preprocessing on training data only and Pandas groupby: retain missing keys and define all-null totals help prevent a tidy metric from hiding a broken data boundary.
Tested environment
Dependency check: this program was executed on CPython 3.14.6 with scikit-learn==1.9.1. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.
Working program
from sklearn.metrics import confusion_matrix, precision_score, recall_score
actual = [0, 0, 0, 1, 1]
predicted = [0, 1, 0, 0, 1]
print(confusion_matrix(actual, predicted, labels=[0, 1]).tolist())
print("precision:", precision_score(actual, predicted, zero_division=0))
print("recall:", recall_score(actual, predicted, zero_division=0))
print("empty-positive precision:", precision_score([0, 0], [0, 0], zero_division=0))Output
[[2, 1], [1, 1]]
precision: 0.5
recall: 0.5
empty-positive precision: 0.0Costs and limits
Counting n labels takes row-dependent work; a k-class confusion matrix retains O(k squared) cells. Accuracy can look high under a skewed class distribution while the rare failure class remains poorly detected.
Common Mistakes
- Declare matrix label order before interpreting its positions.
- Do not treat a zero-division fallback as evidence of adequate positive-class performance.
Connected lessons
Scikit-learn pipelines: fit preprocessing on training data only, Pandas groupby: retain missing keys and define all-null totals, NumPy floating-point checks: finite values and declared tolerances.
