Each one changed a design decision rather than decorating the report.
Document length is the constraint the whole architecture answers to.
Counted against every window a classifier sees, not against other clauses.
Held-out test split. Chunk level asks whether a clause is recognised in the window in front of the model; document level asks whether the contract gets flagged at all.
Ten contracts from the held-out split, chosen to cover all six labels and weighted toward the ones the model gets wrong. Expert annotations are shown in blue; what the model flagged is shown in amber.
Every miss across the ten embedded contracts, by label.
Splits are assigned per contract, never per window, so overlapping windows of one document cannot land on both sides. Decision thresholds are tuned per label on the validation split and applied unchanged to test — tuning them on test would fit the boundary to the answer.
A contract is flagged when its strongest window clears the threshold. The task is existential: a clause in one window of a sixty-window contract means the contract contains it. Averaging instead drops document macro F1 from 0.77 to 0.62.