Ivyra.

How it works

Does your confidence match reality?

You make judgment calls every day — this will ship on time, that hire will work out, I'll stick with the gym this year. Ivyra checks those calls against what actually happens, so you learn where your confidence is trustworthy — and where it runs ahead of the facts. No math required from you.

The problem

We're bad at knowing how much we know

Think of the last time you were sure — "I'm 90% certain we'll close this by Friday." Now think about how often that kind of "90% certain" actually pans out. For most of us, honestly, it's a lot less than 90% of the time.

That gap stays invisible for one reason: we don't keep track. After the fact, memory quietly rewrites itself — "I always knew that was a stretch" — and the lesson evaporates. The only fix is almost embarrassingly simple: write the prediction down, with a number, before you know the answer. That's the whole mechanic here.

Is this just me, or is everyone like this?
It's remarkably universal — decades of research find that people are systematically overconfident, and it barely tracks with intelligence or expertise. The good news buried in that research: it's a trainable skill, not a fixed trait. Weather forecasters, who get scored feedback every single day, end up among the best-calibrated people on earth. Ivyra gives you that same kind of feedback.

The core idea

A probability is a claim about many cases, not one event

This is the one idea everything else rests on, and most people have never had it spelled out. Picture a weather forecaster who says "70% chance of rain." What would make that a good forecast?

Not whether it rains today. A "70% chance" isn't a promise that it will rain — it's a claim that on days like this, it rains about 7 times out of 10. So she's exactly right if, across all the days she says "70%," it rains on about 70% of them. The dry days aren't her being wrong — they're her forecast coming true. Three days in ten, it was supposed to stay dry.

So no single prediction can be "calibrated" or not — only a whole track record can. That's why Ivyra needs a little history before it can tell you much, and why it gets sharper the longer you use it.

Then how can a single prediction be scored at all?
A single prediction still gets an exact score for how close it landed to reality — that's the next section. What it can't tell you is whether your "70%" really means 70%. That only shows up once you've made many of them and we can check: of all your "70%" calls, how often did they come true?

The loop

Four steps, repeated

  1. 1

    Predict

    Write what you think will happen, in your own words. Add a confidence — say, 75% — and the date you'll know.

  2. 2

    Resolve

    When that date arrives, we nudge you. You mark what actually happened: yes or no. Your reasoning stays frozen — no rewriting history.

  3. 3

    Score

    Fixed math — never a guess, never AI — turns that prediction and outcome into an exact number for how close your confidence landed to what happened.

  4. 4

    Adjust

    Do it enough times and the pattern surfaces: where your stated confidence matches how often things actually happen, and where it runs ahead of them. Then you recalibrate.

…and then back to the top, for the next call.

Your score

How close your confidence lands

When a prediction resolves, we measure one thing: how far your confidence was from what happened. Say 90% and it happens — you were barely off. Say 90% and it doesn't — you were badly off. Average that across all your predictions and you get your score.

Lower is better, like golf. A score of 0.25 is the "I'm just guessing" line — exactly what you'd get by shrugging "50/50" at everything. Beat it, and your confidence is carrying real information about the world.

70%

Drag to change how sure you are, then compare the two outcomes below. Lower is better — think of it like golf.

If it happens

Outcome: YES

0.09

Excellent

You were 0.30 away from what happened.

If it doesn't

Outcome: NO

0.49

Costly

You were 0.70 away from what happened.

Notice what happens as you push toward the extremes: being confident and right earns a tiny score, but being confident and wrong is punished far harder than a cautious guess ever is. Sitting at scores exactly 0.25either way — the “I’m just shrugging” baseline. Beating that number, over many predictions, means your confidence actually carries information.

Your bias

Which way you lean

Your score says how good your calls are. Your bias says which direction they're off — the gap between how confident you felt and how often you were actually right.

"You run 19 points overconfident" means that, on average, reality came in about 19 percentage points below your stated confidence. The opposite reading means you were too cautious, and things happened more often than you claimed.

Your curve

The whole picture in one chart

The calibration curve puts your confidence along the bottom and how often things actually happened up the side. If you're spot-on, every dot lands on the diagonal — your 70%s happen about 70% of the time.

Dots below the line mean overconfident: things happened less often than you claimed. Above the line means underconfident — you knew more than you let on. Flip between the shapes below to see it.

0%0%25%25%50%50%75%75%100%100%Your confidenceHow often it happened

Every dot sits below the line: things happen less often than you claimed. When you say 90%, it comes true far less than 90% of the time. This is the common human pattern.

The dashed line is perfect calibration. The gap between each dot and that line is, literally, the size of your self-deception at that confidence level.

Your boldness

Being honest isn't enough — you also have to say something

Here's the part almost everyone misses. There are two separate ways your confidence can fail, and most people only know the first.

One: your numbers can be dishonest. You say 90% and things happen 60% of the time. That's the overconfidence we've been talking about, and the curve catches it.

Two: your numbers can be empty. Imagine answering "60%" to everything. You're never badly wrong — but you're never really saying anything either. You've made yourself safe and useless.

Boldness measures whether your confidence levels actually tell your outcomes apart. Do the things you call 80% really happen more often than the things you call 55%? If they do, your numbers carry information. If everything you say clusters near 50%, they don't — however honest each one looks on its own.

Confidence that's honest and decisive at the same time moves to high and low values as the evidence earns it, and those values match how often things actually happen. Confidence that stays hedged near the middle is safe but says nothing — a hedger's 80% calls land no more often than their 55% ones.

0%50%100%Confidence you gave each call
It happenedIt didn't

Every call hugs 50%. The greens and reds are jumbled together in the middle — the confidence level tells you nothing about which way things went. Never badly wrong, never saying anything.

Reading your results

The verdict and the insight

Your insights page gives you two different things, and it's worth knowing which is which.

The verdict is the one-line summary at the top — "you lean overconfident," "calibrated and bold." It only ever describes what's true about your track record. Fixed math produces it; it states a fact and stops there.

The insight goes further: it explains why, and what to do differently. This is the one place AI helps — it reads your own predictions and the words you wrote, names the pattern behind the numbers, and points to where you can adjust. It never touches your score.

Recent vs. lifetime

Zoom in on how you're doing lately

You can look at your whole history or just your recent calls — and the two can honestly disagree. Your lifetime read might say "overconfident" while your last stretch says "calibrated," because you've been getting better.

That's the point, not a glitch: recent form shows whether the training is working, while lifetime shows the deeper habit. Neither one is the "real" number — they answer different questions.

Honest expectations

Why some things show up later

Because calibration is about patterns, the richer read-outs need a bit of data before they mean anything. A curve built from eight predictions is noise — the same way a coin isn't "biased" for landing heads three times out of four. So Ivyra waits until a number actually means something, rather than guessing early. You get an exact score on your very first resolution; the bigger pictures arrive as you go:

  1. ~10

    Your bias score

    The first real read: a single number like "you run 12 points overconfident," telling you which way your judgment leans and by how much.

  2. ~25

    Your progress chart

    Your score over time, so you can watch whether the training is working — your recent calls against your lifetime average.

  3. ~30

    Your calibration curve and boldness

    The full picture: your confidence against reality across every level. These need the most data, because a curve from a handful of predictions is just noise.

Until each one unlocks, we show you exactly how many resolutions are left — never a blank, never a misleading half-picture.

What you can trust

The math grades you. The AI only explains.

This matters, so we'll say it plainly. Every score, every curve, every number comes from fixed math — the same calculation for everyone, that nothing can nudge. You can't make yourself look better by sounding confident, and neither can we.

AI is used in exactly one way: to read your own words and data and explain the patterns back to you — name the habit, suggest a fix. It never assigns a score, and it only ever sees your own data. The judgment about how good your calls are is pure arithmetic.

Make your first prediction

The whole loop starts with one call about something you actually care about. It takes about thirty seconds.