Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
At a glance
Goal: Test OpenAI API error handling
Time: 10 minutes
Plugins: GenericRandomErrorPlugin, RetryAfterPlugin
Prerequisites: Set up Dev Proxy
When you use OpenAI APIs in your app, you should test how your app handles API errors. Dev Proxy allows you to simulate errors on any OpenAI API using the GenericRandomErrorPlugin. With the RetryAfterPlugin, Dev Proxy also checks that your app waits for the time in the Retry-After header before it calls the API again.
Tip
Download this preset by running in the command prompt devproxy config get openai-throttling.
In your project folder, create a new file named devproxyrc.json. Open the file in a code editor.
Create a new object in the plugins array referencing the GenericRandomErrorPlugin. Define the OpenAI API URL for Dev Proxy to watch and add a reference to the plugin configuration.
File: devproxyrc.json
{
"$schema": "https://raw.githubusercontent.com/dotnet/dev-proxy/main/schemas/v3.3.1/rc.schema.json",
"plugins": [
{
"name": "GenericRandomErrorPlugin",
"enabled": true,
"pluginPath": "~appFolder/plugins/DevProxy.Plugins.dll",
"configSection": "openAIAPI"
}
],
"urlsToWatch": [
"https://api.openai.com/*"
]
}
Add the RetryAfterPlugin and create the plugin configuration object to provide the GenericRandomErrorPlugin with the location of the error responses and the percentage of requests to fail.
File: devproxyrc.json (complete config)
{
"$schema": "https://raw.githubusercontent.com/dotnet/dev-proxy/main/schemas/v3.3.1/rc.schema.json",
"plugins": [
{
"name": "RetryAfterPlugin",
"enabled": true,
"pluginPath": "~appFolder/plugins/DevProxy.Plugins.dll"
},
{
"name": "GenericRandomErrorPlugin",
"enabled": true,
"pluginPath": "~appFolder/plugins/DevProxy.Plugins.dll",
"configSection": "openAIAPI"
}
],
"urlsToWatch": [
"https://api.openai.com/*"
],
"openAIAPI": {
"$schema": "https://raw.githubusercontent.com/dotnet/dev-proxy/main/schemas/v3.3.1/genericrandomerrorplugin.schema.json",
"errorsFile": "openai-errors.json",
"rate": 90
}
}
Caution
Add the RetryAfterPlugin before the GenericRandomErrorPlugin in your configuration file. If you add it after, the GenericRandomErrorPlugin fails the request before the RetryAfterPlugin can check it.
In the same folder, create the openai-errors.json file. This file contains the error responses that Dev Proxy chooses from when it fails a request. They match the errors that the OpenAI API returns:
| Status | error.code |
What it simulates |
|---|---|---|
| 429 | rate_limit_exceeded |
Your app hit its tokens per minute (TPM) or requests per minute (RPM) limit. |
| 429 | slow_down |
Your app's request rate increased too quickly. |
| 429 | credit_balance_exhausted |
Your organization has no prepaid credits left. Retrying doesn't help. |
| 503 | server_is_overloaded |
The model is temporarily overloaded. |
For more information about these errors, see Error codes in the OpenAI documentation.
File: openai-errors.json
{
"$schema": "https://raw.githubusercontent.com/dotnet/dev-proxy/main/schemas/v3.3.1/genericrandomerrorplugin.errorsfile.schema.json",
"errors": [
{
"request": {
"url": "https://api.openai.com/*"
},
"responses": [
{
"statusCode": 429,
"headers": [
{
"name": "content-type",
"value": "application/json; charset=utf-8"
},
{
"name": "Retry-After",
"value": "@dynamic"
}
],
"body": {
"error": {
"message": "Rate limit reached for gpt-4.1 in organization org-K7hT684bLccDbBRnySOoK9f2 on tokens per min (TPM): Limit 30000, Used 30000, Requested 1200. Please try again in 2.4s. Visit https://platform.openai.com/settings/organization/limits to learn more.",
"type": "tokens",
"param": null,
"code": "rate_limit_exceeded"
}
}
},
{
"statusCode": 429,
"headers": [
{
"name": "content-type",
"value": "application/json; charset=utf-8"
},
{
"name": "Retry-After",
"value": "@dynamic"
}
],
"body": {
"error": {
"message": "Rate limit reached for gpt-4.1 in organization org-K7hT684bLccDbBRnySOoK9f2 on requests per min (RPM): Limit 500, Used 500, Requested 1. Please try again in 120ms. Visit https://platform.openai.com/settings/organization/limits to learn more.",
"type": "requests",
"param": null,
"code": "rate_limit_exceeded"
}
}
},
{
"statusCode": 429,
"headers": [
{
"name": "content-type",
"value": "application/json; charset=utf-8"
},
{
"name": "Retry-After",
"value": "@dynamic"
}
],
"body": {
"error": {
"message": "Your request rate increased too quickly. Reduce your request rate and increase it gradually.",
"type": "rate_limit_error",
"param": null,
"code": "slow_down"
}
}
},
{
"statusCode": 429,
"headers": [
{
"name": "content-type",
"value": "application/json; charset=utf-8"
}
],
"body": {
"error": {
"message": "Your organization has no prepaid credits remaining. Add credits to continue using the API. For more information on this error, read the docs: https://developers.openai.com/api/docs/guides/error-codes.",
"type": "insufficient_quota",
"param": null,
"code": "credit_balance_exhausted"
}
}
},
{
"statusCode": 503,
"headers": [
{
"name": "content-type",
"value": "application/json; charset=utf-8"
}
],
"body": {
"error": {
"message": "The requested model is temporarily overloaded. Please try again later.",
"type": "service_unavailable_error",
"param": null,
"code": "server_is_overloaded"
}
}
}
]
}
]
}
The @dynamic value sets the Retry-After header and tells the RetryAfterPlugin to track how long your app must wait. The credit_balance_exhausted response has no Retry-After header, because waiting doesn't fix it.
Start Dev Proxy in your project folder:
devproxy
When your app calls OpenAI APIs, Dev Proxy fails 90% of the requests with a random error from the openai-errors.json file. If your app calls the API again before the time in the Retry-After header, the RetryAfterPlugin reports it and throttles the request.
Check that your app:
- Waits for the
Retry-Aftertime after arate_limit_exceededorslow_downerror. - Stops calling the API after a
credit_balance_exhaustederror, instead of retrying. - Retries with a delay after a
server_is_overloadederror, and shows a clear message when retries run out.
Learn more about the GenericRandomErrorPlugin.