EN ▾

API Proxy FAQ: From Getting a Key to Debugging Errors

This section answers the questions developers ask most often when integrating and using the API, organized into five groups: Getting Started, Quotas and Billing, Errors, Capability Limits, and Content and Privacy. Each answer provides specific numbers and actions so you can quickly find the relevant information for your situation; if you can't find it, return to the documentation page to review the API details.

Updated on

Key points

  1. Register with email and password; your API key appears immediately. New accounts get $0.50 in free trial credit valid for 7 days. No payment info required.
  2. 402 means no balance, 429 means too many requests, and 503 means try again later. Each requires a different response.
  3. Context 100,000 tokens, max output 32,000, request body max 8 MB, text only.
  4. Prompts are not used for training.

Getting Started

How do I get my first API key?

Open the Get Key page, register with email and password, and the key appears immediately after registration. No payment info is needed; copy the key into an environment variable to start calling.

How many keys can one account have?

One per account. You can regenerate it when you need to change it, but note that the old key becomes invalid immediately after regeneration, so all services using the old key will start receiving 401 errors. Prepare the new key distribution process before acting.

What are the endpoint and model name?

The endpoint is https://api.llmzhongzhuan.com/v1,对话请求发往 /v1/chat/completions. View the model list using GET /v1/models. There is only one model name: uncensored.

How do I adapt existing OpenAI code?

Usually, change two things: replace base_url with the address above, replace the key with this site's key, and change the model field to uncensored. The request and response body structure keeps the OpenAI chat completion format, so most code does not need to change. See Framework Configuration for specific framework implementations.

Credits & Billing

How much is the free trial credit and how is it calculated?

New accounts get $0.50, valid for 7 days. At $0.25 per million input tokens and $1.00 per million output tokens, assuming 800 input tokens and 400 output tokens per request, one request costs about $0.0006. The free trial credit can send about 800 such requests. This is an estimate based on assumptions; actual usage depends on your prompt length.

Does the prepaid balance expire?

No. The top-up is prepaid credit. There is no subscription, and no rule for clearing the balance at expiration. The only thing to note is the free trial credit, which expires after 7 days.

How do I know how much a request cost?

The response includes usage, containing input and output token counts. For streaming requests, a chunk with usage is automatically added at the end. Multiply these two numbers by the unit price to get the cost for that request. For monitoring persistence implementations, see Stability Practices.

Where can I see the full pricing?

Pricing is on the Pricing page. It lists separate rates for input and output tokens with no hidden per-request fees.

Errors & Rate Limits

What does a 402 error mean?

Error code no_credit means your balance is empty or the free trial credit has expired. The only solution is to top up prepaid credit. Do not retry this error; retries will fail and only generate extra requests.

How do I handle a 429 error?

Each key allows 300 requests per minute; exceeding this returns 429. Reduce your send rate and implement retry with jittered exponential backoff. For batch processing, use concurrency limits and a rate limiter to smooth out requests. Example code is in the stability article.

Is 503 upstream_busy a permanent failure?

No. It means the service is temporarily busy. Retrying after a few seconds usually succeeds. Use exponential backoff and limit the maximum number of retries.

What causes 403 content_blocked?

The request triggered content filtering. Sexual content involving minors is blocked regardless of whether it is fictional or roleplay. Do not retry on 403; check and modify your request content.

Capabilities

How long is the context, and how long can the output be?

The total context is 100,000 tokens; input and output combined cannot exceed it. max_tokens defaults to 2048 and can be set up to 32,000 per request. If input plus max_tokens exceeds the limit, it returns 400. Leave room in the context for input when writing long texts.

Does it support streaming and function calling?

Yes. Set stream to true for SSE streaming output. Function calling uses the OpenAI tools format. Common sampling parameters like temperature, top_p, and stop are passed through.

Can it do embeddings, images, or audio?

No. This site only provides text chat, with no vector, image, audio, video, or fine-tuning capabilities, and only one model. If your business needs these capabilities, you need to match with other solutions.

Is there a request body size limit?

Yes, a single request body cannot exceed 8 MB. When putting long documents into the prompt, pay attention to the overall byte size in addition to the token count.

Deployment & Operations

What are the most important steps before going live?

First, use /v1/models to confirm the model list, then run tests including long inputs, streaming, and error paths. Then write handling branches for 402, 429, and 503, and write usage to disk. Finally, set a balance alert line to avoid discovering empty credits during peak hours. Do not forget to write these checks into the release process for on-call staff to verify item by item.

What should you check first if all requests suddenly return 401?

First, confirm whether anyone has regenerated the key in the backend, as old keys become invalid immediately after regeneration. Second, check if environment variables were lost during the last deployment, or if spaces and newlines were accidentally included in the key. Keeping the active key in only one place saves most of this troubleshooting effort.

What should you watch out for when multiple services share a single key?

Rate limits are calculated per key, so the total requests from all services must not exceed 300 per minute. If one batch processing task uses up the quota, online services will also receive 429. We recommend limiting the batch processing separately, prioritizing online requests, and staggering batch processing at night if necessary.

How to avoid costs quietly exceeding expectations?

The price per output token is four times that of input tokens, so prioritize controlling max_tokens and the length required in the prompt. Track usage by feature; if the average output for a feature suddenly doubles, check if the prompt was modified. The benefit of the prepaid model is that services stop when the balance is exhausted, preventing unexpected overage bills.

Content and Privacy

What content is allowed?

Legal adult content, fictional works, and controversial topics are not directly rejected; the service is aimed at adult users aged 18 and above. Sexual content involving minors is always blocked. For details on the costs and trade-offs of such services, read Costs and Trade-offs of Uncensored AI API.

Will my prompts be used for training?

No, prompts are not used for training. Aside from this, this page does not provide other storage or logging details; please refer to the documentation.

How does this service differ from an aggregator gateway?

This site offers only one model, so there is no need to map numerous model names. You can verify the model list yourself via /v1/models. To understand the general principles and risks of gateways, read Gateway Principles first.

Frequently Asked Questions

How to get started?

Register with your email and password, and the key will be displayed immediately without requiring payment information. New accounts receive a $0.50 free trial credit valid for 7 days.

What to do if you get a 402?

The 402 error code is no_credit, indicating that the balance is exhausted or the free trial has expired. Top up your prepaid credit; do not retry.

What is the rate limit?

Each key allows 300 requests per minute; exceeding this returns 429. We recommend using a concurrency limit with exponential backoff to smooth out requests.

What is the maximum length for context and output?

The total context window is 100,000 tokens. max_tokens defaults to 2048 and has a maximum of 32,000. The request body must not exceed 8 MB.

Are prompts used for training?

No, prompts are not used for training.

Just fill out the form to get the key

Create an account, copy the key, and modify the Base URL. Configuration is that simple.

Get API key