Jev vs luna-none / luna-low / terra-none / sol-none — 23 subjects / 836 questions, 2026 Common Test (main)
How does TypeSafe AI's Japanese-language performance compare with OpenAI's models?
We chose Japan's university entrance Common Test as a public benchmark that spans Japanese-language knowledge. Questions were sent one at a time — or, when they depended on each other, grouped together as text only — so both models faced identical conditions. Figures and tables were first converted to text by a separate AI.
Subject combinations (normalised to 1000 points):
Pricing: Jev $0.042 per 1M input tokens; luna $0.2 input / $1.2 output. Time is the sum of API latency over all 836 questions.
Every question was passed to the models as text. Figures and tables were transcribed beforehand by a separate AI, and that text was included with the question. Questions are sent one at a time; dependent questions are grouped into a single request. The samples below use dummy content.
POST https://api.typesafe.ai/v1/systemone
{
"model": "jev-latest",
"state": {
"question": "(問題文。図表は文字起こししたテキストをここに含める)",
"figure_note": "(図・表をテキスト化した説明)"
},
"questions": {
"blank_ア": {
"type": "choice",
"instructions": "空欄【ア】に入る最も適当なものを選べ。",
"criteria": { "0": "選択肢0の内容", "1": "選択肢1の内容", "2": "選択肢2の内容" }
}
}
}
POST {OPENAI_URL}/chat/completions
{
"model": "gpt-5.6-luna",
"reasoning_effort": "none",
"messages": [{
"role": "user",
"content": "(問題文+図表のテキスト)\n各空欄の答えを、選択肢キーで答えよ。"
}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "answers",
"strict": true,
"schema": {
"type": "object",
"properties": { "ア": { "type": "string", "enum": ["0", "1", "2"] } },
"required": ["ア"],
"additionalProperties": false
}
}
}
}
reasoning_effort is none (luna-none) or low (luna-low).