DOCUMENTATION

It classifies text into a list of answers you provide. It does not write the answer.

You send text and some possible answers. It returns probabilities.

API · Use-cases · How it's built

API

Base URL https://api.theclassifier.ai. Send Authorization: Bearer nlk_…. Keys are created in the dashboard and shown once.

CLASSIFY

POST https://api.theclassifier.ai/v1/classify
Authorization: Bearer nlk_…
Content-Type: application/json

{
  "text": "Please cancel my account at the end of the month.",
  "choices": ["billing", "technical support", "cancellation", "fraud"]
}
{
  "cancellation": 0.982,
  "billing": 0.011,
  "technical support": 0.004,
  "fraud": 0.003
}

Scores are a relative fit among the choices you sent. They sum to about 1. They are not comparable across different choice lists.

Models are listed on the models page. Pass version to select one. Omit it for the default.

PYTHON

import requests

response = requests.post(
    "https://api.theclassifier.ai/v1/classify",
    headers={"Authorization": "Bearer nlk_..."},
    json={
        "text": "Please cancel my account at the end of the month.",
        "choices": [
            "billing",
            "technical support",
            "cancellation",
            "fraud",
        ],
    },
)
result = response.json()

SEVERAL QUESTIONS

POST https://api.theclassifier.ai/v1/score takes a record and one or more closed questions. Use this when an agent needs more than one decision, or when a choice needs a short criteria string.

{
  "state": "Customer: you charged me twice.",
  "questions": [{
    "question": "What is the customer trying to do?",
    "choices": [
      {"choice": "duplicate_charge", "criteria": "They say they were billed more than once."},
      {"choice": "none", "criteria": "Not a billing problem."}
    ]
  }]
}

Read results[0].selected.choice. If the top two scores are close, ask a clarifying question.

MCP

Streamable HTTP at https://api.theclassifier.ai/mcp. Same API key. The tool is classify.

{
  "mcpServers": {
    "the-classifier": {
      "url": "https://api.theclassifier.ai/mcp",
      "headers": {"Authorization": "Bearer nlk_…"}
    }
  }
}

LIMITS

Two limits apply at once. Tokens are not one of them.

A short classification is about 40 ms and 2,150 input tokens. A long one is about 109 ms and 11,750 input tokens.

If the shared machine is briefly busy, the request waits. If the wait would be longer than the queue limit, the API returns 503 instead of letting latency grow.

When you are over your own limit, the API returns 429:

{
  "error": {
    "type": "rate_limit_exceeded",
    "message": "Compute allowance exceeded.",
    "retry_after_ms": 840,
    "limit_type": "compute"
  }
}

Headers on limited and successful calls:

Retry-After
X-RateLimit-Limit
X-RateLimit-Remaining
X-RateLimit-Reset
X-Compute-Limit
X-Compute-Remaining

BILLING

The monthly fee is a Stripe subscription. That fee is the credits included with the plan. There is no way to add credits.

Successful requests spend those credits. Tokens are summed for the billing period, then priced. A single request is not rounded into a charge.

input_tokens / 1,000,000,000 * A$1

When the included credits are used, the API refuses further calls until the account moves to the next plan:

{
  "error": {
    "type": "credits_exhausted",
    "message": "Credits for this plan are used. Production is required to continue."
  }
}

USE-CASES

The choices are yours. Include none when the text might not fit the list.

Route an agent. The choices are the tools or queues the agent is allowed to call.

Classify a ticket. The choices are your queues: billing, technical support, cancellation, fraud.

Score an LLM response. The choices are the judgments you care about: followed the policy, refused, made up a fact.

Rank a lead. The choices are the bands you already use: hot, warm, cold.

Detect an objection. The choices are the objections you handle, plus none.

Choose between fixed alternatives. Any decision where the answer has to be one of a list you wrote.

HOW IT'S BUILT

The text you send is the record. Each choice is scored for how well it fits that record. A choice can include a short criteria string that says when that choice is true.

The base model is Qwen3 1.7B. Each approved model is a small adapter and a scoring head on that base. The list is on the models page. The default is general.

It does not write, summarize, or invent a label outside the list you sent.

One shared GPU serves the requests. A short call is about 40 ms of GPU time. A long one is about 109 ms. Your plan limits how many requests you can burst and how much of that GPU time you can keep using.

Input tokens are counted so the credits included with the plan can be spent. Tokens do not decide whether the request is allowed. GPU time does.