Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Python classification metrics: fix label order before counting errors

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

A confusion matrix counts actual and predicted labels, exposing different failure types that a single accuracy number can hide.

Download Python source kit

Operation contract

The review fixture declares label order as zero then one. Rows represent actual labels and columns represent predictions. A false positive marks a clean receipt for review; a false negative misses a receipt that required review. Precision and recall use different denominators, so they describe different costs even when their numeric values happen to match.

Failure and ownership boundary

This sample has five fixture records and no claim about real review quality. Empty positive classes require a declared zero-division policy. Threshold selection, class imbalance, grouped records and uncertainty need a larger evaluation design. Scikit-learn pipelines: fit preprocessing on training data only and Pandas groupby: retain missing keys and define all-null totals help prevent a tidy metric from hiding a broken data boundary.

Tested environment

Dependency check: this program was executed on CPython 3.14.6 with scikit-learn==1.9.1. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.

Working program

python
from sklearn.metrics import confusion_matrix, precision_score, recall_score

actual = [0, 0, 0, 1, 1]
predicted = [0, 1, 0, 0, 1]
print(confusion_matrix(actual, predicted, labels=[0, 1]).tolist())
print("precision:", precision_score(actual, predicted, zero_division=0))
print("recall:", recall_score(actual, predicted, zero_division=0))
print("empty-positive precision:", precision_score([0, 0], [0, 0], zero_division=0))

Output

Output
[[2, 1], [1, 1]]
precision: 0.5
recall: 0.5
empty-positive precision: 0.0

Costs and limits

Counting n labels takes row-dependent work; a k-class confusion matrix retains O(k squared) cells. Accuracy can look high under a skewed class distribution while the rare failure class remains poorly detected.

Common Mistakes

  • Declare matrix label order before interpreting its positions.
  • Do not treat a zero-division fallback as evidence of adequate positive-class performance.

Connected lessons

Scikit-learn pipelines: fit preprocessing on training data only, Pandas groupby: retain missing keys and define all-null totals, NumPy floating-point checks: finite values and declared tolerances.

python
classification-metrics
Storage details