OpenKol
Methodology

Here is exactly
what we measure

An audit tool that won't show its working is an opinion with a number attached. Every weight below is the one the scorer actually uses.

What we don't claim
How a check runs

Tweet in, four answers out

Six steps. Everything after the first is deterministic unless you ask for the LLM audit.

1 · Read the post

The post, its author, its timestamp and its replies. Classify intent first — roughly 31% of a KOL's token posts turn out not to be calls at all, and scoring those would poison every average.

2 · Resolve the token

Match the ticker or contract to a live market. Ambiguous tickers stop and ask rather than guessing; a wrong-token score attached to a named person is a dispute waiting to happen.

3 · Pull the chart

Candles either side of the post at four resolutions, plus liquidity and the pair's own pre-post baseline.

4 · Build the benchmark

Sample peer tokens on the same chain over the same window. Returns are scored in excess of that median, so a market-wide melt-up never reads as skill.

5 · Score the replies

Classify each reply by content first and account statistics second. Validated at 91% against hand-labelled posts.

6 · Place it

Compare against the corpus percentiles, assign the profile quadrant, and store the result stamped with the scoring version so an old report stays reproducible.

The four answers

Four scores, never blended

They answer different questions for different people, and one of them has no measurable relationship with the others. Averaging would imply a link the data denies.

performance/60

Was it a good call to follow?

Peak22

Highest excess return over the peer benchmark within the window.

Drawdown13

How far it fell from that peak. A round trip is not a good call.

Hold10

Excess return still standing at the end of the window.

Timing8

How early the entry was relative to the move.

Retention7

Whether the gain persisted rather than spiking and vanishing.

impact/40

Did it move the market?

Volume15

Traded volume above the pair's own pre-post baseline.

Trades11

Distinct trades, so one whale cannot carry the score.

Flow8

Buy pressure against sell pressure in the attribution window.

Immediate6

How much of the move arrived in the first minutes.

audience/40

Is the engagement real?

Comments18

Reply content: substance, relevance, whether anyone engaged with the claim.

Authenticity8

Giveaway entries, wallet drops, contentless praise.

Overlap8

Whether the same accounts reply to every post this KOL makes.

Engagers6

Account-level signals. Weighted low on purpose — see below.

profile

Loud, or actually effective?

Not a score — a label. Engagement per thousand followers and volume per thousand followers, each compared against the corpus p75. The money axis uses raw volume, not the composite impact score, so a big account cannot buy its way into the good quadrant with reach alone.

quiet · no money
quiet · real money
loud · no money
loud · real money
The corpus

What the percentiles are computed from

Scores are relative. A number out of 40 means nothing until you know what the rest of the distribution looks like — so here it is.

Every call
goes in
scored the same way, stored, and re-used as the benchmark
Recalibrated
continuously
thresholds are re-derived as the distribution moves, never hard-coded
Versioned
and reproducible
every stored result carries the version it was computed under
Any chain
we can price
coverage follows the market rather than a fixed list
Measurep25p50p75p90
Impact score5.18.612.818.9
Engagement / 1k0.40.81.42.3
Volume / 1k$60$358$1,200$3,000
Volume lift0.50x0.93x1.58x2.58x

These are measured, and they are lower than intuition suggests. Before calibrating, the engagement threshold was guessed at 12 per thousand followers; the real p75 is 1.4. That single wrong number made two of the four profile quadrants structurally unreachable — which is why nothing here is a fixed constant, and why the thresholds are re-derived as the distribution moves.

What we don't claim

The things that came out badly

These are measurements too. Publishing only the flattering ones would make every other number on this site worth less.

Skill does not persist
Splitting each KOL's history in half and correlating the two halves gives rho −0.024. One account went from +22.8% to −10.7% between its own halves. We therefore never sell a forecast, and there is no “top KOLs to follow” ranking anywhere on this site.
Audience ≠ returns
Audience quality has no measurable relationship with what the token did afterwards: rho 0.052. A real audience is worth knowing about. It is not a predictor, and blending it into one number would pretend otherwise.
Entry timing does predict
The one component with proven predictive value: rho −0.188, p<0.01. Later entries do worse. It is a small effect and we present it as one.
Some checks return 'unknown'
Thin pairs never produce a reliable first-hour volume figure. For those the profile quadrant is withheld rather than guessed. We would rather show nothing than a confident wrong label.
Pump-and-dump is unvalidated
The detector fires on evidence we can point to, but it has zero labelled cases behind it, so its false-positive rate is unknown. It stays private to whoever ran the check, is marked experimental, and always shows its evidence rather than a bare verdict. It will not appear in any public ranking until it has been validated.
The LLM adds nothing here
On the labelled evaluation set the LLM audit scored the same as the deterministic path — 91% both ways. So the free tier runs deterministic, and we don't charge for a model call that didn't improve the answer.
Questions

Details worth asking about

Why is the audience score weighted towards reply content?
Why score returns against peers instead of just the price?
Are trading costs modelled?
What happens with an ambiguous ticker?
Can I reproduce an old report?

Check one post.Stop paying for reach.