Proxy API Stability Practice: Timeouts, Retries, Rate Limit Queues, and Streaming Disconnections
After integrating the proxy API into production, the real drain on your energy is often not the integration itself, but sporadic anomalies: one or two timeouts among dozens of requests, batch processing halted by rate limits, or streaming dropping mid-way. This article handles these systematically from an ops perspective: classify errors first, then provide runnable code for backoff retries, concurrency queues, streaming continuation, and usage monitoring.
Key Points
- Classify before retrying: 429 and 503 warrant backoff retries; 400, 401, 402, 403 retries waste time.
- The 300 requests per minute limit requires two layers: semaphores for concurrency and a rate limiter for speed. One layer is insufficient.
- Set streaming read timeouts as the interval between data blocks. Preserve received content upon interruption and continue writing.
- Persist usage for each request to detect cost anomalies and prompt expansion early.
Classify errors into three categories first
The first step to stability is not adding retries, but knowing what to retry. Common failures can be divided into three categories by handling method:
| Category | Typical Behavior | Handling Method |
|---|---|---|
| Instantly recoverable | 429 rate limit, 503 upstream_busy, connection reset, timeout | Retry with exponential backoff, limiting total attempts |
| Request itself is flawed | 400 (input plus max_tokens exceeds 100,000, request body too large), 403 content_blocked | Do not retry; fix the request or report to the user |
| Account status issue | 401 invalid key, 402 no_credit | Do not retry; alert the owner, top up, or swap the key |
It is best to put this table in code comments. A common incident is: after the balance runs out, the endpoint starts consistently returning 402, and your retry logic does not distinguish status codes, so every request retries five times, needlessly amplifying traffic fivefold and turning the logs red.
Layer timeouts
A single total timeout is insufficient. It is recommended to split timeouts into three layers:
- Connection timeout: Time to establish TCP and TLS connections. 5 to 10 seconds is usually enough; fail fast if it cannot connect.
- Read timeout: time waiting for response data. For non-streaming requests, it must cover the full generation time; for streaming requests, it becomes the maximum silence between two data chunks, typically 10 to 30 seconds.
- Business total timeout: The maximum wait your business can tolerate. Use
asyncio.wait_foror the gateway layer to enforce this, including all retries.
Longer outputs take longer for non-streaming requests. If you set max_tokens to the 32,000 limit but only set a 30-second total timeout, you will frequently timeout on long texts. A safer approach is to always use streaming for long outputs, providing immediate feedback and clarifying timeout semantics. The SDK timeout parameter accepts a number or detailed configuration; check your version's documentation.
Exponential backoff retries for 429 and 503
There are three key points for backoff: delay doubles with each failure, set an upper limit, and add jitter. Without jitter, a batch of tasks failing simultaneously will retry at the same time, crushing the just-recovered service. The function below disables SDK built-in retries, taking unified control ourselves, and retries only on 429, 503, and connection errors:
import asyncio
import os
import random
from openai import AsyncOpenAI, APIConnectionError, APIStatusError, APITimeoutError
client = AsyncOpenAI(
base_url="https://api.llmzhongzhuan.com/v1",
api_key=os.environ["API_KEY"],
timeout=60.0,
max_retries=0, # 重试由下面的函数统一控制
)
RETRY_STATUS = {429, 503}
async def chat_with_retry(messages, max_attempts=5, base_delay=1.0, cap=20.0):
for attempt in range(1, max_attempts + 1):
try:
resp = await client.chat.completions.create(
model="uncensored", messages=messages, max_tokens=800
)
return resp
except APIStatusError as e:
if e.status_code not in RETRY_STATUS or attempt == max_attempts:
raise # 400/401/402/403 重试没有意义
reason = f"HTTP {e.status_code}"
except (APIConnectionError, APITimeoutError) as e:
if attempt == max_attempts:
raise
reason = type(e).__name__
delay = min(cap, base_delay * 2 ** (attempt - 1))
delay = random.uniform(0, delay) # 全抖动,避免所有任务同时醒来
print(f"第 {attempt} 次失败({reason}),{delay:.1f}s 后重试")
await asyncio.sleep(delay)
503 means the server is temporarily busy; retry after a few seconds. The baseline delay of 1 second, cap of 20 seconds, and max 5 attempts cover this scenario well. For 429, if it appears frequently, the issue is not retries but the sender lacking rate control; see the next section.
Concurrency queue under a 300 requests per minute rate limit
Each API key allows 300 requests per minute, which averages to 5 requests per second. Batch processing tasks are most likely to fail here: using asyncio.gather to send thousands of requests at once uses up the quota in the first few dozen, leaving the rest to receive 429.
A reliable approach uses two layers. A semaphore limits the number of concurrent in-flight requests to prevent local connection and memory exhaustion. A rate limiter controls the emission rate to spread requests evenly. Both are essential: with only a semaphore, fast requests can still exceed 300 per minute; with only a rate limiter, slow requests pile up. The example sets the rate to 270 per minute, leaving a 10% buffer for multi-instance, retries, and clock drift:
import asyncio
import os
import time
from openai import AsyncOpenAI
client = AsyncOpenAI(
base_url="https://api.llmzhongzhuan.com/v1",
api_key=os.environ["API_KEY"],
timeout=60.0,
max_retries=0,
)
MAX_IN_FLIGHT = 8 # 同时在途的请求数
PER_MINUTE = 270 # 留 10% 余量,低于每分钟 300 次的上限
class Pacer:
"""把请求的发出时刻均匀铺开:每 60/PER_MINUTE 秒放行一个。"""
def __init__(self, per_minute):
self.interval = 60.0 / per_minute
self.next_at = 0.0
self.lock = asyncio.Lock()
async def wait(self):
async with self.lock:
now = time.monotonic()
delay = self.next_at - now
if delay > 0:
await asyncio.sleep(delay)
self.next_at = max(now, self.next_at) + self.interval
sem = asyncio.Semaphore(MAX_IN_FLIGHT)
pacer = Pacer(PER_MINUTE)
async def summarize(idx, text):
async with sem: # 控制并发
await pacer.wait() # 控制速率
resp = await client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": f"用两句话概括:{text}"}],
max_tokens=120,
)
return idx, resp.choices[0].message.content
async def main():
texts = [f"第 {i} 条工单的正文……" for i in range(1000)]
tasks = [summarize(i, t) for i, t in enumerate(texts)]
done = 0
for coro in asyncio.as_completed(tasks):
idx, out = await coro
done += 1
if done % 100 == 0:
print(f"已完成 {done} 条")
asyncio.run(main())
If multiple processes or machines share one key, the in-memory rate limiter is insufficient because limits are merged by key. You need a shared token bucket, commonly implemented with Redis counting, or simply route all batch requests through a dedicated queue consumer service that handles the egress.
What to do when streaming disconnects
Streaming connections are more fragile than regular requests: they last longer, traverse more network devices, and any idle timeout can cut them off. The principle is that received content is an asset; do not discard it all just because of an interruption.
The following implementation accumulates received chunks. Upon catching connection or timeout exceptions, it sends the existing text as an assistant message to request continuation. This is a best-effort recovery; the continuation may not be seamless, making it suitable for long texts and conversations but not for strict structures like JSON. If structured output is interrupted, restarting entirely is safer. Note that streaming automatically attaches a usage chunk at the end; the code extracts it simultaneously:
import asyncio
import os
from openai import AsyncOpenAI, APIConnectionError, APITimeoutError
client = AsyncOpenAI(
base_url="https://api.llmzhongzhuan.com/v1",
api_key=os.environ["API_KEY"],
timeout=30.0, # 流式场景下,这是相邻数据块之间允许的最长静默
max_retries=0,
)
async def stream_once(messages):
parts, usage = [], None
stream = await client.chat.completions.create(
model="uncensored", messages=messages, max_tokens=600, stream=True
)
try:
async for chunk in stream:
if chunk.usage: # 最后一个分片携带 usage
usage = chunk.usage
if chunk.choices and chunk.choices[0].delta.content:
parts.append(chunk.choices[0].delta.content)
except (APIConnectionError, APITimeoutError):
return "".join(parts), None, False # 中断:返回已收到的部分
return "".join(parts), usage, True
async def stream_with_resume(question, max_resume=2):
messages = [{"role": "user", "content": question}]
text = ""
for _ in range(max_resume + 1):
part, usage, finished = await stream_once(messages)
text += part
if finished:
return text, usage
# 把已生成部分作为 assistant 消息,请模型接着写
messages = [
{"role": "user", "content": question},
{"role": "assistant", "content": text},
{"role": "user", "content": "请从上一条回复的结尾处继续,不要重复已写的内容。"},
]
return text, None
if __name__ == "__main__":
out, usage = asyncio.run(stream_with_resume("用三段话说明为什么日志要带 request id。"))
print(out)
print("usage:", usage)
When continuing, remember that sent text counts as input tokens, so each continuation incurs extra cost. Limiting the number of max_resume attempts is necessary.
Degradation, circuit breaking, and pre-launch checklist
Retries handle transient faults. If a fault persists for minutes, continuing to retry only piles up requests. Add a circuit breaker layer above retries: if consecutive failures reach a threshold (e.g., ten failures in one minute), pause sending requests for a short period, routing directly to degradation logic (e.g., returning cached results, prompting the user to try later, or requeuing tasks). After the cooldown, allow only a few probe requests; restore full volume if they succeed.
Degradation must also be planned. For user-facing chat, return a friendly busy message; for offline batch tasks, write failed entries to a failure table and retry them collectively after the run finishes, rather than waiting indefinitely in the main flow. In either case, ensure failures are not silently swallowed; at least log a line with the request id.
Finally, here is a pre-launch checklist: Did you disable implicit SDK retries and keep only your own? Did you give up on retrying 400, 401, 402, 403? Does the total timeout include all retries? Does batch processing have rate control instead of one-time concurrency? Is streaming interruption handled? Is usage persisted? Is there an alert for balance? Pass these items one by one before load testing, not the other way around.
Monitoring and reconciliation using usage
Every response includes usage, appearing in the last chunk for streaming. Persisting it along with the function name is the cheapest monitoring method. Our input price is $0.25 per million tokens and output is $1.00; you can estimate the cost of each request locally:
import csv
import time
LOG = "usage_log.csv"
def record_usage(feature, resp):
u = resp.usage
cost = u.prompt_tokens * 0.25 / 1_000_000 + u.completion_tokens * 1.00 / 1_000_000
with open(LOG, "a", newline="", encoding="utf-8") as f:
csv.writer(f).writerow(
[int(time.time()), feature, u.prompt_tokens, u.completion_tokens, f"{cost:.6f}"]
)
After deployment, check three metrics daily: the median input token count per function to detect prompt bloat; the output token ratio, since output pricing is four times the input rate and usually dominates the bill; and the 429/503 ratio, where an increase signals that your send rate or retry strategy needs adjustment. Add a balance alert to notify the responsible person to top up before a 402 error occurs. For more integration details, see the framework configuration, and common issues are summarized in the FAQ.
FAQ
Should you retry immediately on 429?
No. Back off with random jitter and check if the sender exceeds 300 requests per minute. Persistent 429s indicate a need for a rate limiter, not more retries.
How long to wait for 503 upstream_busy?
A few seconds. Start with 1-second exponential backoff, cap at ~20 seconds, limit total attempts, and return a clear failure to the caller if exceeded.
If a streaming request breaks, should you discard received content?
No. Keep the received text and request a continuation as an assistant message; for strict structures like JSON, a full retry is safer.
How to calculate the cost per request?
Read usage from the response: multiply input tokens by 0.25 USD per million, output tokens by 1.00 USD per million. For streaming, usage is in the last chunk.
How is the rate limit calculated when multiple machines share one API key?
It is aggregated per key; total requests across all machines must not exceed 300 per minute. Use shared counting or a unified egress queue to coordinate.
Fill out the form to get your API key
Create an account, copy the key, and update the Base URL. Configuration is that simple.