← Back to Directory

Inference.net: Schematron V2 Turbo

INFERENCE-NET Developer Architecture Profile

Intelligence (ELO)1039Chatbot Arena Verified
Max Context Limit128,000Max Out: 8,192 tokens
Prompt CachingStandard$0.03 / 1M cached
Standard API / 1M$0.18Blended Standard Rate

Model Capabilities & Contract SLAs

    Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

    Granular Pricing Matrix

    Input Tokens (Standard Prompt)$0.03 / 1M
    ⚡ Cached Input (Prompt Cache Read)% OFF$0.03 / 1M
    Output Tokens (Completion)$0.15 / 1M

    Pricing data via OpenRouter. Sync: 10/2/2026

    Evaluate Competitors