What is an API Gateway: Principles, Common Risks, and a Selection Checklist
Many developers first encounter an “API gateway” and know only that changing the address and key lets them call an LLM, but cannot explain what happens in between. This article dissects the full request path from an ops perspective, clarifying forwarding, keys, and billing, then lists the three most common pitfalls and a selection checklist, ending with two commands to verify it yourself.
Key points
- The essence of a gateway is “proxy forwarding + key mapping + usage accounting”. Your request takes one extra hop, and stability and security depend on that hop.
- Three most common pitfalls: improper key storage, receiving a model different from what you expected, and rate limits defined outside the documentation.
- Don't just look at unit price; check if the model list is visible, error codes are standard, and quota and rate limits are clearly stated.
- After getting the key, run /v1/models and a small request once. You can rule out most problems within ten minutes.
The path a request takes through a gateway
First, let's clarify the terminology. An "API Proxy" is a gateway placed between your code and the backend model provider, exposing a standard interface. Your code still sends requests in OpenAI format, but you point base_url to the gateway and use the key it provides.
The gateway typically does three things in this hop.
- Request forwarding: Validates the request body, fills in default parameters if necessary, and passes the request to the backend. The backend’s response (including streaming SSE chunks) is returned to you as-is or with light processing.
- Key mapping: You hold a key issued by the gateway, which is only meaningful within the gateway. The gateway uses it to identify you, your balance, and which models you can call. The credentials actually used with the backend remain internal to the gateway and never appear in your code.
- Billing & Rate Limits: After each request, the gateway deducts balance based on input/output tokens multiplied by the unit price, and counts requests per key per minute, returning 429 if exceeded.
Linking these three things together explains why gateway experiences vary greatly: the forwarding layer determines latency jitter and streaming stability, the key layer determines the scope of loss if leaked, and the billing layer determines if bills are transparent and auditable.
Difference from direct connection to a single model
Direct connection means sending requests directly to the official domain of the model provider. Usually, one account corresponds to one set of models, one billing rule, and one set of documentation. Proxy services typically come in two forms, differing by what is connected behind them.
| Dimension | Direct single service | Aggregated gateway | Single-model gateway |
|---|---|---|---|
| Number of models | A few from the provider | Dozens or even hundreds | One |
| Interface format | Proprietary formats | Unified to OpenAI-compatible format | OpenAI compatible |
| Troubleshooting difficulty | Lowest, shortest chain | Highest, many model name mappings | Lower, only one model |
| Suitable scenarios | Stable business using only one provider | Frequent model switching for comparison | Fixed model, seeking predictability |
If your business relies on only one model, the benefits of aggregation are useless, and you bear the uncertainty of “which model the name maps to”. Conversely, if you switch models weekly for comparison tests, an aggregated service saves a lot of adaptation work. There is no absolute superiority; the key is knowing which category you belong to.
This site falls into the last category: it offers only one model with the ID uncensored and an OpenAI-compatible chat completion interface. For a discussion on these trade-offs and costs, see Unlimited AI API: Cost & Trade-offs.
Three most common risks
Key security
A gateway key is equivalent to a prepaid credit card: whoever gets it can spend your balance. Common leak paths include: writing the key into frontend code, committing it to public repositories, or pasting it into tickets or group chat screenshots. It is recommended to place it only in server-side environment variables, and have the frontend always route through your own backend. If you suspect a leak, reset immediately; the old key should become invalid instantly. Also check if the service allows self-service reset and whether the old key is invalidated immediately after reset, rather than “taking effect after a few hours”.
Model substitution
This is the most discussed issue in aggregated services: you request A, but get back cheaper B. It's hard to judge from docs; you must verify via behavior. Fix a set of questions with standard answers, set temperature, and test repeatedly to see if output style is stable. You can also request /v1/models to check if the list matches the pricing page. Vague model names or inconsistent behavior over time are warning signs.
Opaque rate limits
Some services only write “reasonable use” in the documentation, but secretly slow down or drop requests during peak hours, causing your program to show occasional timeouts. A mature approach states the number of requests per minute for each key, returns a standard 429 when exceeded, rather than leaving connections hanging. When selecting, be sure to ask: are rate limits per key or per account, what is returned when exceeded, and is there a separate error code when the balance is exhausted?
Checklist for evaluating proxy services
Copy this list directly into your evaluation document and tick each item off.
- Does it expose a public
GET /v1/modelsendpoint that returns a model list matching the pricing page? - Are error responses structured JSON containing code and message, with distinct codes for 401, 402, 429, and 503?
- Is the per-key request-per-minute rate limit documented, rather than just mentioned by support?
- Are the context window length, maximum output tokens per request, and request body size specified with clear numbers?
- Is billing deducted precisely by the token count in the usage field, and can the balance be viewed at any time?
- Does the prepaid credit expire? Is the validity period of free trial credit clearly stated?
- Can keys be reset by yourself? Do old keys become invalid immediately?
- Does it support streaming output, and does the final response include usage statistics for your own reconciliation?
- Is there a clear one-sentence statement regarding whether prompts are used for training?
- Are unsupported capabilities (e.g., vectors, images, audio) honestly marked, rather than vaguely described?
A perfect score is unrealistic, but if you cannot answer the first five items, we recommend trying with a small amount first rather than topping up a large balance at once.
Ten-minute verification after obtaining the key
Regardless of which provider you choose, it is worth spending ten minutes on basic verification before going live. Step one: list the models and confirm that the returned IDs match your expectations:
curl -s https://api.llmzhongzhuan.com/v1/models \
-H "Authorization: Bearer $API_KEY"
Step two: send a small request and observe whether the usage field exists in the response and whether the numbers are reasonable. The example below forces the model to recite a date to observe whether it hallucinates information it cannot know. This is a rough behavioral check, not a rigorous evaluation:
curl -s https://api.llmzhongzhuan.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "用一句话介绍你自己,然后复述今天的日期是几号。"}],
"max_tokens": 200
}'
Add these two steps to your deployment script and run them whenever you change keys or services. If the response lacks usage data, or if the token count doesn't match the input length, billing transparency is poor, so clarify before spending large amounts. To learn how to integrate into various frameworks, see Framework Configuration Guide.
Our parameters for your reference against the checklist
Here we list our actual parameters so you can check them against the checklist above without flipping back and forth through documentation.
- Endpoint:
https://api.llmzhongzhuan.com/v1, supportingPOST /v1/chat/completionsandGET /v1/models, authenticated with a Bearer key. - Only one model, id
uncensored; text only, no vectors, images, audio, video, or fine-tuning. - 100,000 token context window (input plus output),
max_tokensdefaults to 2048 with a maximum of 32,000 per request; request body size does not exceed 8 MB. - Each key is limited to 300 requests per minute, returning 429 when exceeded; 503 upstream_busy indicates you should retry later; 402 no_credit is returned when the balance is exhausted or the trial expires.
- Pricing is $0.25 per million input tokens and $1.00 per million output tokens. Prepaid credit, no subscription, balance never expires.
- Prompts are not used for training.
Specific numbers are subject to the pricing page and documentation. New accounts receive a $0.50 free trial credit valid for 7 days. No payment information is required to register, so you can use it to complete the verification process above.
Frequently asked questions
What is the biggest difference between an API Proxy and calling the official API directly?
The proxy adds a gateway layer between you and the model, responsible for forwarding, key rotation, and billing. The longer chain buys you a unified interface format and more flexible billing, at the cost of requiring you to trust this layer's stability and integrity.
How to tell if a proxy service is secretly swapping models?
Test repeatedly with fixed questions and fixed temperature to see if the output is stable, and verify that the /v1/models list matches the pricing page. Single-model services have fewer uncertainties because there is only one id.
What if a proxy key is leaked?
Reset the key immediately in the backend and confirm whether the old key becomes invalid instantly. In the future, keep keys only in server-side environment variables and have the frontend route through your own backend.
What should you check first when choosing a proxy service?
First check whether the model list is publicly queryable, whether error codes are standardized, and whether the per-minute rate limit is documented. Then check the context window length and whether the balance expires. Compare unit prices after these.
How much credit is safe to test with first?
Use the free trial credit or a small balance to run /v1/models and a few typical requests, then gradually increase usage. We do not recommend topping up a large amount at the start.
Just fill out the form to get the key
Create an account, copy the key, and modify the Base URL. Configuration is that simple.