Configuring the API Gateway in Common Frameworks: From SDK to Dify
Connecting to a relay API requires only three values: the endpoint, the API key, and the model name. The difficulty lies in the fact that different frameworks use different names for these three values: some call it base_url, others api_base, and Dify uses a form. This article lists them side by side, every code snippet is ready to run, and a troubleshooting checklist is provided at the end.
Key Points
- Put the endpoint and API key into environment variables API_BASE and API_KEY to avoid hardcoding.
- The endpoint must include the /v1 suffix, and the model name is fixed as uncensored.
- LangChain uses base_url, while LlamaIndex's OpenAILike uses api_base. The names differ, but the meaning is the same.
- In Dify, select the 'OpenAI-API-compatible' provider and manually enter the model name, endpoint, and context window length.
Prepare the three values first
Regardless of the framework, confirm you have these three items first. The rest is just filling in the blanks.
- Endpoint:
https://api.llmzhongzhuan.com/v1. Keep the trailing/v1, but do not add/chat/completions; the SDK handles that. - API Key: Displayed immediately after registration with email and password on the Get Key page. Each account has one key. You can reset it, and the old one becomes invalid immediately.
- Model Name: There is only one,
uncensored. You can verify this yourself usingGET /v1/models.
Also note these limits: the context window is 100,000 tokens, the default max_tokens is 2048 (max 32,000), and each API key allows 300 requests per minute. These numbers will be relevant when configuring parameters.
Environment Variable Syntax
Hardcoding the API key is a common source of errors. The recommended approach is to read only from environment variables, injected by your container or secret management system during deployment. The examples below use API_KEY and API_BASE.
# Linux / macOS:写进 ~/.zshrc 或 .env 加载脚本
export API_KEY="替换为你的密钥"
export API_BASE="https://api.llmzhongzhuan.com/v1"
# Windows PowerShell(仅当前会话)
$env:API_KEY = "替换为你的密钥"
$env:API_BASE = "https://api.llmzhongzhuan.com/v1"
If you use a .env file, remember to add it to .gitignore. Also, the official Python SDK reads OPENAI_API_KEY and OPENAI_BASE_URL by default. You can use these names or pass them explicitly in the constructor. Explicit parameters prevent interference from leftover variables on your machine, simplifying debugging.
OpenAI Python SDK and Node SDK
Python requires v1+ of the openai package; Node requires v4+. The difference is only syntax. Pass the endpoint and API key when constructing the client. First, a non-streaming Python call that prints usage for billing verification:
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ.get("API_BASE", "https://api.llmzhongzhuan.com/v1"),
api_key=os.environ["API_KEY"],
)
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "写一条 20 字以内的发布公告标题。"}],
max_tokens=100,
)
print(resp.choices[0].message.content)
print(resp.usage)
The Node example uses streaming, which is the standard way to implement the typewriter effect in the frontend. The server automatically appends a final chunk with usage when the streaming request ends, so no extra parameters are needed:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.API_BASE ?? "https://api.llmzhongzhuan.com/v1",
apiKey: process.env.API_KEY,
});
const stream = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "列出三个命名服务端日志字段的好习惯。" }],
stream: true,
max_tokens: 300,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");
The Node example uses top-level await, requiring a .mjs file or setting "type": "module" in package.json. If your runtime does not support this, wrap it in an async function.
LangChain: base_url for ChatOpenAI
LangChain does not require any special adapters. Simply use ChatOpenAI from the langchain_openai package and point base_url to your endpoint.
import os
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="uncensored",
base_url=os.environ.get("API_BASE", "https://api.llmzhongzhuan.com/v1"),
api_key=os.environ["API_KEY"],
temperature=0.7,
max_tokens=500,
timeout=60,
)
print(llm.invoke("把“服务已降级”改写成对用户友好的一句话。").content)
Two tips. First, explicitly set max_tokens and timeout, as defaults may not suit your use case. Second, if you use function calling in your chain, we support OpenAI-format tools, so LangChain's bind_tools works normally. For streaming, consume chunks inside stream().
LlamaIndex: OpenAILike
LlamaIndex's built-in OpenAI class validates if the model name is in the official list and errors on custom names. Use OpenAILike from llama-index-llms-openai-like instead. It skips validation and uses different parameter names, such as api_base for the endpoint.
import os
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
model="uncensored",
api_base=os.environ.get("API_BASE", "https://api.llmzhongzhuan.com/v1"),
api_key=os.environ["API_KEY"],
is_chat_model=True,
context_window=100000,
max_tokens=500,
)
print(llm.complete("用一句话解释什么是幂等请求。"))
is_chat_model=True is critical here; it forces LlamaIndex to use the chat endpoint instead of the legacy completion endpoint. Set context_window to 100000 so LlamaIndex does not truncate your documents based on the default few thousand tokens.
Dify: OpenAI-API-compatible Provider
Dify is a visual platform configured via forms rather than code. Interface text may vary slightly by version, but the steps are consistent:
- Go to 'Settings', find the 'Model Provider' page, select 'OpenAI-API-compatible' from the list, and click 'Add Model'.
- Select 'LLM' as the model type and enter
uncensoredas the model name. - Enter your API Key in the API Key field and
https://api.llmzhongzhuan.com/v1in the API endpoint URL field. - Set the model context window to 100000 and the max tokens limit to 32000.
- If you need function calling in your workflow, enable the function calling support option. Keep streaming enabled.
- Save, create a simple chat app, select the newly added model, and send a message to verify it returns correctly.
The platform may send a probe request upon saving. If this fails, it is usually due to an extra path in the endpoint or extra spaces in the API key. If your Dify is deployed in a container, ensure the container can access external domains.
Parameters and Deployment Checks Before Launch
Getting the framework to run is just the first step. Verify these parameters before launch, as they determine cost and failure rates.
- max_tokens: Default is 2048. Increase it explicitly for long outputs, up to a max of 32,000. Note that input plus output cannot exceed 100,000 tokens, or you will receive a 400 error.
- timeout: Long outputs take more time. In streaming scenarios, set the read timeout to over 60 seconds. For non-streaming, estimate based on the longest output.
- temperature / top_p / stop: These standard sampling parameters are passed through. Values set in the framework take effect directly without extra toggles.
- Parallel requests: 300 requests per minute per key. When multiple service instances share the same key, the rate limit is calculated collectively, so do not plan for 300 requests per instance.
- Retries: Framework retries often only cover network errors. Add backoff logic for 429 and 503 errors yourself. See the stability guide for implementation.
When deploying to containers, the principle remains the same: do not bake the API key into the image. Inject it via your orchestration tool at startup. Here is the minimal compose syntax:
# docker-compose.yml 片段
services:
app:
image: your-app:latest
environment:
API_BASE: https://api.llmzhongzhuan.com/v1
API_KEY: ${API_KEY} # 从宿主机环境或 .env 读取,不写进镜像
In Kubernetes, switch to a Secret reference; the principle remains the same. Additionally, it is recommended to prepare different accounts or API keys for different environments, such as one set for development, staging, and production. This way, if an API key leaks or runs out of balance in one environment, it will not affect the online environment. Since each account has only one API key, it is best to correspond multiple environments to multiple accounts and pre-top up an appropriate balance for each.
Finally, check your logs. Many frameworks print full request headers in debug mode, including Authorization. Disable this debug output in production or mask the API key in your log filters. Recording the usage field is also good practice; it helps you reconcile bills and spot if a feature's prompt suddenly grows longer.
Troubleshooting Order for Errors
90% of integration issues cluster in a few areas. Check in this order for the fastest resolution:
- 401: The key is empty, contains extra spaces when copied, or you are still using the old key after resetting it.
- 404: The base URL points to the root path without
/v1, or you appended/chat/completionstwice. - 402: The error code is no_credit, meaning your balance is exhausted or the free trial credit has expired. You need to top up your prepaid credit.
- 400: This is usually caused by the input plus
max_tokensexceeding 100,000, or the request body exceeding 8 MB. - 429 / 503: The former is a rate limit of 300 requests per minute; the latter is upstream_busy. Wait a few seconds and retry. See stability practices for details.
Another often-overlooked issue is the network environment. Corporate intranets, proxy software, or security group rules may allow domain name resolution but fail to establish an HTTPS connection, resulting in long hangs followed by timeouts rather than explicit error codes. In such cases, first try accessing /v1/models with curl on the same machine; if it works, the issue is at the application layer; if not, check the proxy and firewall. Document the troubleshooting process in team docs so new colleagues can avoid this detour next time.
If all three values are correct but the request still fails, use curl to make a direct request first to rule out framework-specific issues. If you want to understand what the proxy does under the hood, refer to how the proxy works.
Frequently asked questions
Does base_url need to include /v1?
Yes. Use https://api.llmzhongzhuan.com/v1. The SDK will automatically append /chat/completions, so do not add it manually.
Why does LlamaIndex use OpenAILike instead of OpenAI?
The OpenAI class checks if the model name is in the official list; custom model names will cause an error. OpenAILike performs no validation, making it better suited for compatible interfaces.
Can I name the model arbitrarily in Dify?
No, the model name must be uncensored because the request carries this ID as-is. The display name can be set separately, but the model field must match.
Why don't changes to environment variables take effect?
Usually, the terminal or process has not been restarted. We recommend passing base_url and api_key explicitly in your code and printing the actual address used to confirm.
Fill out the form to get your API key
Create an account, copy your key, and modify the Base URL. Configuration is that simple.