Quickly integrate an uncensored large model API
Connect an uncensored large model to your app in minutes with basic configuration and simple code examples. We provide an OpenAI-compatible API proxy service that supports streaming and tool calling, helping you quickly build apps without content limits.
Basic Configuration
Connecting to the API proxy core requires correctly configuring the Base URL and authentication credentials. Our service is fully compatible with the OpenAI protocol. You only need to set the Base URL to https://api.llmzhongzhuan.com/v1 and include your API key in the request headers.
Each account is assigned only one API key, which is displayed immediately upon registration. If you need to change it, you can regenerate it in the backend; the old key becomes invalid immediately. This design simplifies the authentication process and ensures access security. We do not support parallel management of multiple keys; we recommend managing keys reasonably according to your app architecture.
Generative Chat Interface
The standard chat interface is located at POST /v1/chat/completions. The request body must include the model (fixed to uncensored), the messages array, and optional parameters. This interface supports non-streaming responses, suitable for scenarios requiring complete return results.
The model runs on independently deployed high-performance GPU servers with a context window of up to 100,000 tokens. Whether processing long texts or following complex instructions, it provides a stable native API access experience.
curl https://api.llmzhongzhuan.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Streaming Response (SSE)
For apps requiring real-time feedback, enable the stream: true parameter to get chunked responses via Server-Sent Events (SSE). Streaming significantly reduces Time To First Token (TTFT), improving user experience.
The server continuously pushes JSON fragments until the response ends. The client must assemble these fragments to generate the complete text. This mode is ideal for chatbots or real-time content generation scenarios, ensuring data reaches end users quickly.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmzhongzhuan.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Model List Query
You can get the list of currently available models via GET /v1/models. The request returns a standard JSON response containing fields like id, object, and owned_by.
Currently, there is only one model ID: uncensored. This is an open-weight model, specifically fine-tuned to remove content filters, suitable for adult content, safety research, and controversial topics. Note that we do not provide Embedding or multimodal generation services; we focus on pure text chat.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmzhongzhuan.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Error Handling and Rate Limiting
The service follows standard HTTP status codes. 401 Unauthorized means the API Key is invalid; 402 Payment Required means the account balance is insufficient and requires a top up before you can make further requests; 429 Too Many Requests means you have hit the rate limit, which is currently set to 300 requests per minute.
Additionally, the single request body size is limited to 8 MB. If you encounter rate limiting, we recommend implementing an exponential backoff retry strategy. Our uncensored AI service ensures fair resource allocation while maintaining high throughput.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Example Code
Below is a complete example of making a streaming request using the Python SDK. The code shows how to initialize the client, build message history, and handle SSE data streams.
Make sure you have installed the openai library and replace YOUR_API_KEY with your actual key. This example intuitively shows how to get real-time generation results from the API proxy interface, making it easy for developers to integrate into existing systems.
Capabilities and limits
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Item | Value |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Model ID | uncensored |
| Authentication | Bearer token in the Authorization header |
| Base URL | https://api.llmzhongzhuan.com/v1 |
| Completion length | up to the rest of the 100,000-token window; max_tokens optional (no separate cap) |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Context window | 100,000 tokens (prompt + completion together) |
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| JSON mode | JSON object mode via response_format json_object |
| Concurrency | 8 requests at the same time per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Rate limit | 300 requests per minute per key |
| Max body | up to 8 MB per request |
| Free trial | $0.50 for 7 days, no card · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Subscription | no monthly fee; paid credit does not expire |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Keys | one active key per account; a new key replaces the old one |
| Sign-in | sign in with Google or with e-mail + password |
| Content | adult content allowed; sexual content involving minors is refused |
When a request fails
Every error is JSON with a type you can switch on. You are never charged for an error.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | unknown endpoint |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Frequently Asked Questions
Can an account have multiple API keys?
Each account is strictly limited to one API key. To change it, generate a new key in the backend; the old key becomes invalid immediately. This design simplifies permission management, ensuring unique and clear access credentials for each account.
Can an API key be used for multiple apps?
Yes, a single API key can be reused across multiple apps or environments as long as they share the same account quota. Since each account has only one key, we recommend managing multi-app usage by monitoring request frequency and balance.
Does the uncensored model include all adult content?
The model has zero content filters for lawful adult content, fictional stories, and controversial topics, but strictly prohibits sexual content involving minors. This is the only hard content limit, designed to meet basic compliance requirements. Other types of adult content are generated normally.
Just fill out the form to get the key
Create an account, copy the key, and modify the Base URL. Configuration is that simple.