Common Test AI comparison

Jev vs luna-none / luna-low / terra-none / sol-none — 23 subjects / 836 questions, 2026 Common Test (main)

日本語

Purpose

How does TypeSafe AI's Japanese-language performance compare with OpenAI's models?

Method and why

We chose Japan's university entrance Common Test as a public benchmark that spans Japanese-language knowledge. Questions were sent one at a time — or, when they depended on each other, grouped together as text only — so both models faced identical conditions. Figures and tables were first converted to text by a separate AI.

Conclusion

  1. Jev is both smarter and faster than luna-none (no chain-of-thought).
  2. luna-low (with chain-of-thought) beats Jev on accuracy, but Jev is overwhelmingly cheaper and faster.
  3. Among the reasoning-off models, sol-none scores highest and rivals luna-low, while terra-none falls below luna-none.

Data

At a glance

Subject combinations (normalised to 1000 points):

Scores by subject

Scores by question

Category
Subject

Speed & cost

Pricing: Jev $0.042 per 1M input tokens; luna $0.2 input / $1.2 output. Time is the sum of API latency over all 836 questions.

Download

Subject JSON
Questions (source): Questions come from the 2026 Common Test (main) published by Japan's National Center for University Entrance Examinations. NCTU — 2026 Common Test (main)

Request format

Every question was passed to the models as text. Figures and tables were transcribed beforehand by a separate AI, and that text was included with the question. Questions are sent one at a time; dependent questions are grouped into a single request. The samples below use dummy content.

Jev — TypeSafe System One (Choice)

POST https://api.typesafe.ai/v1/systemone

{
  "model": "jev-latest",
  "state": {
    "question": "(問題文。図表は文字起こししたテキストをここに含める)",
    "figure_note": "(図・表をテキスト化した説明)"
  },
  "questions": {
    "blank_ア": {
      "type": "choice",
      "instructions": "空欄【ア】に入る最も適当なものを選べ。",
      "criteria": { "0": "選択肢0の内容", "1": "選択肢1の内容", "2": "選択肢2の内容" }
    }
  }
}

luna — OpenAI-compatible (structured output)

POST {OPENAI_URL}/chat/completions

{
  "model": "gpt-5.6-luna",
  "reasoning_effort": "none",
  "messages": [{
    "role": "user",
    "content": "(問題文+図表のテキスト)\n各空欄の答えを、選択肢キーで答えよ。"
  }],
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "answers",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": { "ア": { "type": "string", "enum": ["0", "1", "2"] } },
        "required": ["ア"],
        "additionalProperties": false
      }
    }
  }
}

reasoning_effort is none (luna-none) or low (luna-low).