A raw 70 percent from a model often behaves like 55 percent after a few hundred labeled trades. MarketXED’s learning loop tracks what happened after a pick and remaps displayed confidence toward frequencies you actually observed. The point is to be less wrong about being wrong, especially for size. This is not a promise of edge and not financial advice.
Why raw scores lie
Committee conviction is an internal number: votes, agreement, and a conservative prior for the scanned universe. On gainer and small-cap tapes, continuation-style hits have historically been modest. If the UI printed “80 percent” from ensemble enthusiasm, you would size as if the next ten trades were almost free. Calibration exists so that a printed probability is closer to a hit rate than to a vibe.
The projector already bounds raw ensemble odds in a humble range and shrinks a thin sample toward a default prior. Calibration sits on top of that honesty. It does not invent a new market thesis; it asks whether last month’s 62s paid like 62s.
How labels are written
Outcomes are labeled from practical paths: did price tag the stop, the target, or a time stop first? That is closer to a triple-barrier idea than to a next-close guess. A timer exit that was slightly green is not the same event as a target hit. Mixing those labels is how dashboards fake skill. MarketXED keeps action picks and monitor noise in separate buckets so the learner is not trained on every scan flicker.
- Hit: the path you cared about completed.
- Miss: stop or invalidation won.
- Flat / time: the clock ended the experiment.
Those labels feed isotonic-style mapping and an online update of agent weights. Agents that were anti-correlated in this universe can be flipped. Near-random seats shrink toward silence.
What you should do with a calibrated number
Use it to decide pass versus take, then let the playbook translate risk dollars into shares, stop distance, and a hold window. Do not treat 58 percent as “almost 60, so double size.” Half-Kelly logic in the playbook already assumes the probability is not a trophy. If the loop has few settled action picks, expect the gate to stay strict (minimum conviction still applies) and the printed odds to stay conservative.
The committee article explains where the raw score came from. Calibration is the second sentence: after the vote, after the label, after enough n. Until n is honest, the correct SEO-friendly claim is simply that MarketXED refuses to perform uncalibrated confidence theater.
What calibration cannot do
It cannot learn multi-day squeezes if the playbook exits on a 90-minute timer. It cannot fix a universe that is all already-extended gainers. It cannot waive PDT or cash settlement. If you change hold rules, you need new labels; old probabilities belong to the old experiment. That is the whole scientific point of a learning loop.
Sample size and the conviction gate
Isotonic maps need repeats. A handful of settled action picks will not move the curve much, and a noisy week should not swing the universe prior. That is why a shrink constant exists and why minimum conviction can stay strict while n is small. If the UI looks “quiet,” that can be the honest state of a young scoreboard, not a broken scanner.
Separate experiments if you change hold time, stop multiple, or universe. Mixing overnight hopes into a 90-minute label file is how people conclude the calibrator is “wrong.” It is labeling a different game. Publish the rule you traded, then let the loop see only that rule. Until then, read probabilities as capped, conservative, and provisional.
Worked example: timer green versus target
Entry 100, stop 98.50, target 102.40, clock 90 minutes. At minute 90 the print is 100.40. That is a timer exit, slightly green. If you label it HIT the same as a target tag, the isotonic map will think 58 percent names pay like targets. They do not. Mark time or flat, keep the HIT bucket for the barrier you advertised. Tomorrow the gate stays honest.
If n of settled action picks is under a few dozen, treat every printed probability as a prior plus a hint. Do not build a size schedule from three lucky timers. The shrink constant is doing you a favor. Your job is to keep feeding it clean labels, not to override it because a week felt smart.