Model Efficiency
Watch impactLive published free itemWhat Are AI Tokens?
A token is the unit a model actually bills and limits you on. Not a word, not a character. This is the foundation piece: read it before Burn Radar, before Burn, before any of it.
What changed
A token is a chunk of text, roughly three to four English characters or about three quarters of a word. "Tokenomics" is unhappy shorthand, not a currency. When you send a prompt, every word in it, every file you paste, every tool definition your harness loads silently, and every word the model writes back all count as tokens. Providers charge per token on metered access, and cap you in tokens (however the vendor labels the cap) on a subscription or a company seat. Context window, the amount a model can "see" at once, is also measured in tokens, not messages or minutes.
Why it matters
Nearly every efficiency question on this desk reduces to the same shape: something is consuming more tokens than the useful work justifies. A repo map resent on every turn. A tool definition nobody calls. A model two tiers too expensive for a one-line classification. None of that is visible until you know tokens are the actual unit being spent, not "usage" as a vague feeling. Burn, as we define it elsewhere on the desk, is tokens consumed relative to useful output delivered. You cannot judge that ratio without first knowing what is being counted.
Take
A token is the unit a model actually bills and limits you on. Not a word, not a character. This is the foundation piece: read it before Burn Radar, before Burn, before any of it.
What to do
Next time you hit a rate limit, a context warning, or a bill you did not expect, ask what actually got tokenised: the visible prompt, or everything riding along with it (history, tool definitions, retries, pasted files). That question is where every other guide on this desk starts.
Evidence
No public sources are linked yet.