Get API key

LLM Leaderboard: Uncensored LLM API Quickstart Guide

This quickstart guide walks you through connecting to our uncensored LLM API, covering authentication, basic requests, streaming, and tool calling with concrete examples.

Base URL & Authentication

Our API follows the OpenAI compatibility standard, meaning you can use existing SDKs with minimal configuration. The base URL is https://api.uncensoredllmleaderboard.com/v1. To authenticate, generate an API key from your dashboard and pass it in the Authorization header as a Bearer token. Each account gets one key, which you can regenerate at any time to revoke access. Unlike many llm api provider options, we do not require a credit card for the initial trial, and prompts are not used for training. Remember that our model, identified as "uncensored", is an open-weight model tuned to answer without refusals for lawful adult content, distinct from GPT or Claude.

Chat Completions Endpoint

The core functionality is text-in, text-out via the POST /v1/chat/completions endpoint. You send a message array and receive a generated response. This endpoint supports streaming via Server-Sent Events (SSE) and tool calling. Our model has a 100,000 token context window for both prompt and completion. Unlike aggregated leaderboards that only show scores, we provide immediate developer access to this uncensored chat completions endpoint with predictable, upfront pricing. There are no embeddings, image, or audio capabilities here—just raw text generation. Use the following example to send your first request:

curl https://api.uncensoredllmleaderboard.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK

For Python developers, the official OpenAI SDK works out of the box. You need to set the base_url and api_key to point to our infrastructure. This approach avoids writing custom HTTP clients and leverages familiar methods like client.chat.completions.create(). The SDK handles JSON serialization and retry logic automatically. Ensure you are using a recent version of the openai package. The model ID you must specify is "uncensored". Here is how to initialize the client and make a request:

from openai import OpenAI

client = OpenAI(base_url="https://api.uncensoredllmleaderboard.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node SDK

Node.js developers can use the official openai npm package. Similar to Python, you configure the client with our base URL and your API key. This allows you to integrate the uncensored model into your existing Node applications without rewriting request logic. The SDK supports both synchronous and asynchronous patterns. Remember that our service is a llm api provider focused on raw token costs, not enterprise SLAs. Use this pattern to interact with the chat completions endpoint:

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.uncensoredllmleaderboard.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

For lower latency and real-time user experiences, enable streaming by setting stream: true. The API returns a stream of Server-Sent Events (SSE), allowing you to process tokens as they are generated. This is useful for chat interfaces or applications where waiting for the full response is too slow. The SDKs handle the parsing of SSE chunks automatically. Note that streaming does not change the pricing or limits; you still pay for the total tokens used. Here is how to implement streaming in your code:

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Rate Limits & Quotas

Our API enforces a limit of 300 requests per minute per key and an 8 MB request body size. If you exceed the rate limit, the API returns a 429 status code. Authentication failures result in a 401 error, typically due to an invalid or missing key. If your prepaid credit runs out, you will receive a 402 error. Credits are paid via crypto (USDT or USDC), starting at $10, with bonuses for larger top-ups. Unlike some open ai api pricing models, we use a simple pay-as-you-go prepaid credit system with no monthly fees. The trial credit of $0.50 is valid for 7 days and requires no card. Context length is capped at 100k tokens, so monitor your input size to avoid exceeding limits or incurring higher costs.

Specs at a glance

If your tool speaks the OpenAI API, these are the details that matter.

SpecValue
ProtocolOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
Model IDuncensored
MethodsPOST /v1/chat/completions · GET /v1/models
API keyAuthorization: Bearer YOUR_KEY
Base URLhttps://api.uncensoredllmleaderboard.com/v1
Max context100,000 tokens, input and output combined
Structured outputJSON object mode via response_format json_object
SSE streamingYes — server-sent events; the last chunk carries token usage
Function callingYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Max outputup to 16,000 tokens per request (default 2,048)
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Rate limit300/min per key
Parallel requests8 requests at the same time per key
Request sizeup to 8 MB per request
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Billingpay as you go from prepaid credit; nothing is charged for failed or refused requests
Trial credit$0.50 of credit valid 7 days, no card needed
Subscriptionno monthly fee; paid credit does not expire
Token prices$0.25 per 1M input tokens · $1.00 per 1M output tokens
Top-upcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Volume bonus+5% on $50+, +10% on $100+
Contentuncensored for adults; the only hard rule: no sexual content involving minors
AccountGoogle or e-mail and password
Keysone key per account, regenerate any time (the old one stops working)

HTTP errors

Every error is JSON with a type you can switch on. You are never charged for an error.

CodeTypeMeaning
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busytemporary overload, retry shortly

Questions and answers

Is this the same as GPT or Claude?

No. Our model is an open-weight model identified as "uncensored". It is not GPT, Claude, Gemini, or DeepSeek. It is tuned to answer without refusals for lawful adult content and is served from our own GPU servers.

How does pricing work?

We charge $0.25 per 1M input tokens and $1.00 per 1M output tokens. You pay via prepaid credit, which never expires. There are no subscriptions or monthly fees. Credits can be topped up starting at $10.

Do you store my prompts?

We do not use your prompts for training. An account requires only an email and password. We do not offer SLA guarantees or certifications like SOC2, but we prioritize privacy for uncensored chat completions.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key