How it works
Does your confidence match reality?
You make judgment calls every day — this will ship on time, that hire will work out, I'll stick with the gym this year. Ivyra checks those calls against what actually happens, so you learn where your confidence is trustworthy — and where it runs ahead of the facts. No math required from you.
The problem
We're bad at knowing how much we know
Think of the last time you were sure — "I'm 90% certain we'll close this by Friday." Now think about how often that kind of "90% certain" actually pans out. For most of us, honestly, it's a lot less than 90% of the time.
That gap stays invisible for one reason: we don't keep track. After the fact, memory quietly rewrites itself — "I always knew that was a stretch" — and the lesson evaporates. The only fix is almost embarrassingly simple: write the prediction down, with a number, before you know the answer. That's the whole mechanic here.
Is this just me, or is everyone like this?→
The core idea
A probability is a claim about many cases, not one event
This is the one idea everything else rests on, and most people have never had it spelled out. Picture a weather forecaster who says "70% chance of rain." What would make that a good forecast?
Not whether it rains today. A "70% chance" isn't a promise that it will rain — it's a claim that on days like this, it rains about 7 times out of 10. So she's exactly right if, across all the days she says "70%," it rains on about 70% of them. The dry days aren't her being wrong — they're her forecast coming true. Three days in ten, it was supposed to stay dry.
So no single prediction can be "calibrated" or not — only a whole track record can. That's why Ivyra needs a little history before it can tell you much, and why it gets sharper the longer you use it.
Then how can a single prediction be scored at all?→
The loop
Four steps, repeated
- 1
Predict
Write what you think will happen, in your own words. Add a confidence — say, 75% — and the date you'll know.
- 2
Resolve
When that date arrives, we nudge you. You mark what actually happened: yes or no. Your reasoning stays frozen — no rewriting history.
- 3
Score
Fixed math — never a guess, never AI — turns that prediction and outcome into an exact number for how close your confidence landed to what happened.
- 4
Adjust
Do it enough times and the pattern surfaces: where your stated confidence matches how often things actually happen, and where it runs ahead of them. Then you recalibrate.
…and then back to the top, for the next call.
Your score
How close your confidence lands
When a prediction resolves, we measure one thing: how far your confidence was from what happened. Say 90% and it happens — you were barely off. Say 90% and it doesn't — you were badly off. Average that across all your predictions and you get your score.
Lower is better, like golf. A score of 0.25 is the "I'm just guessing" line — exactly what you'd get by shrugging "50/50" at everything. Beat it, and your confidence is carrying real information about the world.
Drag to change how sure you are, then compare the two outcomes below. Lower is better — think of it like golf.
If it happens
Outcome: YES
0.09
Excellent
You were 0.30 away from what happened.
If it doesn't
Outcome: NO
0.49
Costly
You were 0.70 away from what happened.
Notice what happens as you push toward the extremes: being confident and right earns a tiny score, but being confident and wrong is punished far harder than a cautious guess ever is. Sitting at scores exactly 0.25either way — the “I’m just shrugging” baseline. Beating that number, over many predictions, means your confidence actually carries information.
Your bias
Which way you lean
Your score says how good your calls are. Your bias says which direction they're off — the gap between how confident you felt and how often you were actually right.
"You run 19 points overconfident" means that, on average, reality came in about 19 percentage points below your stated confidence. The opposite reading means you were too cautious, and things happened more often than you claimed.
Your curve
The whole picture in one chart
The calibration curve puts your confidence along the bottom and how often things actually happened up the side. If you're spot-on, every dot lands on the diagonal — your 70%s happen about 70% of the time.
Dots below the line mean overconfident: things happened less often than you claimed. Above the line means underconfident — you knew more than you let on. Flip between the shapes below to see it.
Every dot sits below the line: things happen less often than you claimed. When you say 90%, it comes true far less than 90% of the time. This is the common human pattern.
The dashed line is perfect calibration. The gap between each dot and that line is, literally, the size of your self-deception at that confidence level.
Your boldness
Being honest isn't enough — you also have to say something
Here's the part almost everyone misses. There are two separate ways your confidence can fail, and most people only know the first.
One: your numbers can be dishonest. You say 90% and things happen 60% of the time. That's the overconfidence we've been talking about, and the curve catches it.
Two: your numbers can be empty. Imagine answering "60%" to everything. You're never badly wrong — but you're never really saying anything either. You've made yourself safe and useless.
Boldness measures whether your confidence levels actually tell your outcomes apart. Do the things you call 80% really happen more often than the things you call 55%? If they do, your numbers carry information. If everything you say clusters near 50%, they don't — however honest each one looks on its own.
Confidence that's honest and decisive at the same time moves to high and low values as the evidence earns it, and those values match how often things actually happen. Confidence that stays hedged near the middle is safe but says nothing — a hedger's 80% calls land no more often than their 55% ones.
Every call hugs 50%. The greens and reds are jumbled together in the middle — the confidence level tells you nothing about which way things went. Never badly wrong, never saying anything.
Reading your results
The verdict and the insight
Your insights page gives you two different things, and it's worth knowing which is which.
The verdict is the one-line summary at the top — "you lean overconfident," "calibrated and bold." It only ever describes what's true about your track record. Fixed math produces it; it states a fact and stops there.
The insight goes further: it explains why, and what to do differently. This is the one place AI helps — it reads your own predictions and the words you wrote, names the pattern behind the numbers, and points to where you can adjust. It never touches your score.
Recent vs. lifetime
Zoom in on how you're doing lately
You can look at your whole history or just your recent calls — and the two can honestly disagree. Your lifetime read might say "overconfident" while your last stretch says "calibrated," because you've been getting better.
That's the point, not a glitch: recent form shows whether the training is working, while lifetime shows the deeper habit. Neither one is the "real" number — they answer different questions.
Honest expectations
Why some things show up later
Because calibration is about patterns, the richer read-outs need a bit of data before they mean anything. A curve built from eight predictions is noise — the same way a coin isn't "biased" for landing heads three times out of four. So Ivyra waits until a number actually means something, rather than guessing early. You get an exact score on your very first resolution; the bigger pictures arrive as you go:
- ~10
Your bias score
The first real read: a single number like "you run 12 points overconfident," telling you which way your judgment leans and by how much.
- ~25
Your progress chart
Your score over time, so you can watch whether the training is working — your recent calls against your lifetime average.
- ~30
Your calibration curve and boldness
The full picture: your confidence against reality across every level. These need the most data, because a curve from a handful of predictions is just noise.
Until each one unlocks, we show you exactly how many resolutions are left — never a blank, never a misleading half-picture.
What you can trust
The math grades you. The AI only explains.
This matters, so we'll say it plainly. Every score, every curve, every number comes from fixed math — the same calculation for everyone, that nothing can nudge. You can't make yourself look better by sounding confident, and neither can we.
AI is used in exactly one way: to read your own words and data and explain the patterns back to you — name the habit, suggest a fix. It never assigns a score, and it only ever sees your own data. The judgment about how good your calls are is pure arithmetic.
Make your first prediction
The whole loop starts with one call about something you actually care about. It takes about thirty seconds.