Why should we use TOON instead of JSON when sending data to LLM models, and how does it reduce token usage?
05:07 15 Nov 2025

I’m trying to understand the advantages of using TOON (Token-Optimized Object Notation) instead of traditional JSON when sending structured data to LLM models such as GPT or Claude.

Many prompt-engineering examples recommend TOON because it supposedly reduces token count and therefore lowers cost per API call. However, I’m not fully clear on:

  1. What exactly is TOON and how is it structured?

  2. How is TOON different from JSON in a way that LLMs interpret with fewer tokens?

  3. How does TOON reduce token usage per request in real LLMs?

  4. What are the trade-offs—such as readability, tooling, or strictness?

Below is an example of the same data in JSON vs. TOON:

JSON

{
  "user": {
    "name": "Alex",
    "age": 27,
    "preferences": {
      "theme": "dark",
      "language": "en"
    }
  }
}

TOON

user:
  name: Alex
  age: 27
  preferences:
    theme: dark
    language: en

From what I understand, TOON saves tokens because it removes quotes, braces, commas, and other punctuation. But I want to confirm whether LLMs genuinely tokenize this format more compactly, and whether there are measurable savings.

If anyone can explain the reasoning or provide benchmark comparisons, that would be very helpful.

json openapi large-language-model slm-phi3