CUAD v1 · multi-label clause classification

Clause Risk Review

A TextCNN reads commercial contracts and flags six risk-relevant clause types. This page reports what the corpus actually looks like, how the model performs against it, and — on held-out contracts it has never seen — where it is wrong.

Four findings that shaped the build

Each one changed a design decision rather than decorating the report.

Contracts are far longer than the models that read them

Document length is the constraint the whole architecture answers to.

The real class balance

Counted against every window a classifier sees, not against other clauses.

Positive rate per label

Labels co-occur

Contracts containing both labels. Off-diagonal mass is why the head is six independent sigmoids rather than a softmax over exclusive classes.

How the model performs

Held-out test split. Chunk level asks whether a clause is recognised in the window in front of the model; document level asks whether the contract gets flagged at all.

Per-class F1

chunk level document level

Training

Loss falls throughout; validation F1 plateaus, and the checkpoint kept is the best epoch rather than the last.

Per-class detail

Contract explorer

Ten contracts from the held-out split, chosen to cover all six labels and weighted toward the ones the model gets wrong. Expert annotations are shown in blue; what the model flagged is shown in amber.

Contract
Document scores against each label's threshold
Clause text

Where it fails

Every miss across the ten embedded contracts, by label.

How the numbers were produced

Splitting and thresholds

Splits are assigned per contract, never per window, so overlapping windows of one document cannot land on both sides. Decision thresholds are tuned per label on the validation split and applied unchanged to test — tuning them on test would fit the boundary to the answer.

Pooling

A contract is flagged when its strongest window clears the threshold. The task is existential: a clause in one window of a sixty-window contract means the contract contains it. Averaging instead drops document macro F1 from 0.77 to 0.62.