Quickstart for Apertus users
Get started with the uncensored API in minutes. This quickstart covers the base URL, authentication, and the chat completions endpoint using your preferred language.
Base URL & Auth
Configure your client to point at the Apertus base URL. Authentication uses a simple API key passed in the Authorization header. You generate this key during signup via Google or email. The key remains static until you manually rotate it. Keep it secure, as it controls your prepaid credit.
curl https://api.apertus.top/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'The base URL is https://api.apertus.top/v1. This ensures compatibility with any OpenAI-compatible SDK. You do not need to modify the SDK initialization logic beyond swapping the base URL and the API key.
First Request
Send a standard chat completion request. The endpoint accepts POST requests to /v1/chat/completions. Provide the model ID and your message content. The response returns the generated text and token usage.
from openai import OpenAI
client = OpenAI(base_url="https://api.apertus.top/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)This request uses the uncensored model. It is an open-weight model tuned for fewer refusals. It is not GPT, Claude, or any other vendor's model. The API returns a JSON object with the assistant's reply.
Model ID
Always use the model ID uncensored in your requests. This ID corresponds to the single large language model hosted on our servers. It is not a routing service. You are interacting directly with this specific uncensored LLM.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.apertus.top/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Do not attempt to switch models. The API does not support multiple model choices. The model is tuned to answer without content refusals for lawful adult use. It supports a context window of 64,000 tokens.
Streaming Responses
Enable streaming by setting stream: true. The API returns a Server-Sent Events (SSE) stream. Each chunk contains partial text. The final chunk includes the full token usage statistics.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Streaming is useful for real-time user interfaces. It does not affect pricing. You are charged for the total tokens processed, not the number of chunks. Errors in streaming are handled by the client SDK.
Function Calling
Supports function calling via the tools parameter. Define your functions in the request body. The model can return tool calls in JSON format.
Use the tool_choice parameter to control when the model uses tools. This feature works with the standard OpenAI schema. Ensure your client parses the tool calls correctly. The model is optimized for accurate tool usage without unnecessary refusals.
Limits, Errors & Context
The context window is 64,000 tokens total. The maximum output is 16,000 tokens, or 2,048 if max_tokens is unset. Rate limits are 300 requests per minute and 8 concurrent requests per key. The request body must not exceed 8 MB.
Common errors include 401 for an invalid key, 402 if prepaid credit is exhausted, and 429 for rate limits. Errors do not consume tokens. Credit never expires. Use JSON mode for structured outputs by setting response_format to {"type": "json_object"}.
Questions and answers
Is this an official site for another vendor?
No. Apertus is an independent service. We do not serve other vendors' models. We provide a single dedicated uncensored model accessible via an OpenAI-compatible interface.
How is billing handled?
Billing is prepaid via crypto. You top up with USDT (TRC20) or USDC (Base). Credit is charged per token used. Errors and refusals are free. Credit never expires.
What is the maximum output length?
The maximum output is 16,000 tokens per request. If you do not set <code>max_tokens</code>, the limit defaults to 2,048 tokens. The total context window is 64,000 tokens.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.