Introduction

In 2026, Codex CLI has matured past the basic installation phase. The main engineering challenge now lies in redirecting the client from official OpenAI endpoints to self-hosted inference services, third-party model providers, and internal gateways. Each custom endpoint requires a properly written config.toml file to negotiate parameters such as base_url, model identifiers and authentication keys. A typo in the model field or misconfigured provider section can lead to immediate request failures. This guide compiles hands-on configuration practices, common error patterns and resolution steps from production deployment experience.

This material targets DevOps engineers and platform builders preparing to integrate Codex CLI into internal gateways, independent developers using third-party OpenAI-compatible endpoints to cut inference costs, and new users stuck at config.toml loading failures after installing Codex CLI.

1. Why Teams Migrate to Compatible Interfaces

1.1 What Codex CLI Is

Codex CLI is OpenAI’s command-line coding agent. It operates inside terminal environments and functions as an AI programming assistant. Given a development task, it can read project files, search source code, execute shell commands, evaluate runtime results, and iterate until task completion. It works in conversational mode instead of processing isolated one-shot prompts.

The default configuration is tightly coupled with official OpenAI services: it points to api.openai.com and authenticates via ChatGPT accounts or OpenAI API keys. Many developers want to route Codex CLI toward local model deployments, private fine-tuned LLMs or self-hosted stacks. Not every organization relies on OpenAI for coding workloads, and strict data residency rules often prohibit sending proprietary code outside the internal network. Connecting Codex CLI with OpenAI-compatible gateways became a mainstream engineering practice throughout 2025 and 2026.

1.2 Value and Typical Scenarios of Compatible Interfaces

An OpenAI-compatible service exposes an HTTP API surface that matches OpenAI’s official specification, including request paths, JSON schema and response structure. When the server implements this standard, clients built for OpenAI can be migrated with minimal configuration changes. Codex CLI, which natively targets OpenAI-style APIs, can connect to nearly any LLM service supporting this standard via configuration adjustments. The common use cases fall into three categories.

  • Internal Enterprise LLM Gateways: Platforms such as OneAPI, NewAPI and LiteLLM manage credentials and rate limits across multiple model vendors, exposing a unified entry point for internal teams.
  • Third-party Model Services: Providers including DeepSeek, Moonshot, Qwen and Ollama offer OpenAI-compatible endpoints. Developers can connect Codex CLI without rewriting client logic.
  • Local Private Deployment: Engineers spin up local inference stacks using Ollama or vLLM. With a compatible API layer, Codex CLI can run entirely offline for sensitive code workloads, with no outbound traffic to public cloud APIs.

In all these scenarios, config.toml acts as the single control file. Mastering this file allows engineers to connect Codex CLI to any OpenAI-shaped service endpoint.

2. Understanding config.toml: Location, Structure and Loading Rules

2.1 Configuration File Location

The default path for the configuration file is ~/.codex/config.toml. If the environment variable CODEX_HOME is defined, the configuration directory changes to $CODEX_HOME/config.toml. Codex CLI does not read .config/.codex automatically, and users cannot override the configuration path through simple command-line window count parameters. This detail is frequently overlooked during initial setup.

The fastest method to verify whether the configuration loads correctly is the debug command:

codex --debug

This command prints the loaded configuration path and active model profile. The log line Loaded config from /Users/xxx/.codex/config.toml confirms successful loading.

2.2 Configuration Loading Logic

config.toml is not a flat dictionary; it contains multiple distinct blocks: global parameters including model and model_provider, [model_providers.xxx] provider definitions, [hooks], and [experimental]. Codex CLI parses the whole file on startup. A syntax error in any field can cause full parsing failure.

A critical behavior: the top-level model field depends on model_provider for resolution. When Codex reads model = "gpt-4o", it does not assume this is an official OpenAI model. Instead, it looks up the provider defined by model_provider inside the [model_providers] table and forwards requests through that provider. The value of model only names the model recognized by the backend service; Codex CLI does not validate whether the model exists at OpenAI. This explains why changing model alone often has no effect — model_provider controls which endpoint the client uses.

3. Line-by-Line Breakdown of a Working config.toml

3.1 Starting from the Official Default Profile

When Codex CLI runs for the first time, it generates a default config.toml under ~/.codex/. File structure varies between releases, and the core template looks like this:

model = "gpt-5-codex"
model_provider = "openai"

When model_provider = "openai", Codex uses the built-in OpenAI provider definition, which is not written explicitly into the TOML file. Its implicit equivalent is:

[model_providers.openai]
name = "OpenAI"
base_url = "[https://api.openai.com/v1](https://api.openai.com/v1)"
env_key = "OPENAI_API_KEY"
wire_api = "responses"

These five lines illustrate the core mechanism: the top-level model_provider selects an entry point, and the [model_providers.xxx] block defines the endpoint properties.

3.2 Core Configuration Fields for Compatible Endpoints

When connecting to an OpenAI-compatible service, focus on these key fields:

  • name: Human-readable provider label, does not affect API requests.
  • base_url: API entry address, ending at /v1. Do not append /chat/completions or /responses. Codex appends suffixes automatically based on wire_api.
  • env_key: Environment variable name storing the API key. If set to MY_LLM_KEY, Codex reads $MY_LLM_KEY from shell environment.
  • wire_api: Protocol type, accepts responses or chat. Official OpenAI endpoints use responses, while most third-party compatible services implement the chat completions protocol. Selecting the wrong value triggers format mismatch errors.

Below is an example connecting a third-party gateway exposing OpenAI-compatible APIs:

model = "deepseek-chat"
model_provider = "third-party"

[model_providers.third-party]
name = "YOUR API Gateway"
base_url = "[https://treerouter.com/v1](https://treerouter.com/v1)"
env_key = "YOUR_API_KEY"
wire_api = "chat"

In this configuration, model must match the model identifier recognized by the backend gateway. After saving the file, define the environment variable inside the shell:

export YOUR_API_KEY="sk-xxxxx"

3.3 Authentication: API Key vs Login Session

Codex CLI supports two authentication modes: native ChatGPT account login, and API key authentication.

For custom providers with env_key defined, Codex prioritizes the value from the specified environment variable as the Authorization Bearer token. In this case, codex login is unnecessary. If env_key is unset but base_url exists, Codex attempts to reuse an existing login session token for the custom endpoint. This behavior creates confusion inside private gateways that do not accept OpenAI tokens. Best practice: define env_key inside your custom provider block, so Codex reads keys from independent environment variables and bypasses official login state.

3.4 Full Annotated Configuration Example

The following complete configuration works in production environments, with inline comments for modification reference.

# Global settings: select model and provider
model = "qwen-max-latest"
model_provider = "company-gateway"

# Hook: execute custom scripts on session start
[hooks]
[hooks.SessionStart]
command = ["echo", "Session starts at $(date)"]

[model_providers.internal]
name = "Internal Gateway"
base_url = "[http://10.20.30.40:8080/v1](http://10.20.30.40:8080/v1)"
env_key = "INTERNAL_LLM_KEY"
wire_api = "chat"

Two frequent pitfalls appear here. If base_url is written as [https://llm-gateway.corp.example.com/v1/chat/completions](https://llm-gateway.corp.example.com/v1/chat/completions), Codex appends another path segment automatically, resulting in duplicated paths such as /v1/chat/completions/chat/completions and returning 404 errors. In addition, query_params support varies by Codex CLI version and may fail parsing on older releases.

4. Hands-On Deployment: Point Codex CLI Toward Custom Endpoints

4.1 Installation and Environment Setup

The standard installation method uses npm:

npm install -g @openai/codex

Installation failures usually stem from network issues. Switch npm registry to resolve connectivity problems:

npm config set registry [https://registry.npmmirror.com](https://registry.npmmirror.com)

Validate the installed version after setup:

codex --version

The first launch runs interactive prompts and asks for ChatGPT account login if no custom provider exists.

4.2 Write and Load the Custom Configuration

Suppose an internal gateway is running at [http://10.20.30.40:8080/v1](http://10.20.30.40:8080/v1), implementing OpenAI chat completions with model name internal-llm-70b. The config.toml entry would be:

model = "internal-llm-70b"
model_provider = "internal"

[model_providers.internal]
name = "Internal Inference Gateway"
base_url = "[http://10.20.30.40:8080/v1](http://10.20.30.40:8080/v1)"
env_key = "INTERNAL_LLM_KEY"
wire_api = "chat"

Export the environment variable:

export INTERNAL_LLM_KEY="sk-xxxx"

If the configuration is valid, Codex skips the ChatGPT login prompt and enters a ready state. Run the simple status command for validation:

/status

When the output shows Model: internal-llm-70b, the connection succeeds.

4.3 Validate Endpoint Before Troubleshooting Codex

After writing the configuration, do not immediately debug Codex CLI. Test the raw endpoint with curl first to confirm the service itself responds:

curl [http://10.20.30.40:8080/v1/chat/completions](http://10.20.30.40:8080/v1/chat/completions) \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-xxxx" \
-d '{
"model": "internal-llm-70b",
"messages": [{"role":"user", "content":"say ok"}],
"max_tokens":20
}'

A normal JSON response confirms the backend service works. If curl fails, the issue lies with the gateway service rather than Codex configuration.

One subtle detail: Codex chat protocol requests include tools and tool_choice fields. Even if curl tests pass, the client may report "tools unavailable" when the gateway or model lacks tool calling support. This is a common trap for OpenAI-compatible integrations.

4.4 Multi-Provider Switching Strategy

Engineers often need to toggle between different providers. Codex CLI supports runtime model switching inside chat sessions, but changing to a fully separate provider requires reloading the environment.

A clean workflow on Linux and macOS uses CODEX_HOME environment variables to point toward separate configuration directories. Each directory holds its own independent config.toml.

CODEX_HOME=~/.codex-internal codex
CODEX_HOME=~/.codex-thirdparty codex

This approach avoids repeated edits to a single configuration file and keeps provider profiles isolated. For teams managing multiple model backends through a unified routing plane, Treerouter, an API gateway, can centralize model access, token consumption tracking and routing rules, reducing repeated configuration work for Codex CLI across different environments.

5. Troubleshooting Common Error Patterns

5.1 Configuration File Parsing Failures

Typical error messages: Failed to load config, cannot load config.toml. These errors usually come from TOML syntax mistakes.

  • Unquoted string values: model = gpt-5 is invalid; use model = "gpt-5".
  • Duplicate keys within one table block.
  • Typo in field names: baseurl instead of base_url.
  • Provider names containing spaces or special characters without quotation marks.

Debugging workflow: run codex --debug. The log points to the line and column number of parsing errors.

5.2 Authentication and Network Errors

Common codes: AuthenticationError, 401 Unauthorized, 403 Forbidden, connection timeout.
Checklist:

  1. Confirm the environment variable defined in env_key exists.
  2. Verify the API key value is correct.
  3. Check whether the gateway requires custom request headers. Custom static headers can be added inside provider definition:
[model_providers.internal.http_headers]
X-API-ID = "secret-value"

Keep in mind that plaintext secrets inside TOML files carry security risks. Set file permission chmod 600 to restrict read access.

Timeouts and SSL errors require inspection of base_url protocol and certificates. Self-signed HTTPS certificates cannot bypass validation in Codex CLI; use plain HTTP or trusted certificates for private endpoints.

5.3 Model and Protocol Mismatch

Errors include model xxx not found, unexpected response format, unsupported wire_api.

  1. Confirm the model ID matches the identifier recognized by your backend gateway. The name written in model must exactly match the service side definition.
  2. Check wire_api. If your endpoint implements /chat/completions but wire_api = "responses", requests return 404 or malformed responses. Change wire_api to chat.
  3. Some compatible endpoints only implement partial OpenAI schema. If responses lack choices[0].message.tool_calls, Codex cannot invoke tools, even if basic chat works.

5.4 Terminal and File Tool Unavailable

Error message: No terminal or file tools available. This indicates Codex wants to execute shell commands or read project files, but tool capabilities are disabled.
Three root causes:

  1. The underlying model does not support tool calling. Codex relies on model-native function calling definitions to invoke tools.
  2. Security restrictions block terminal execution in sandbox or restricted environments.
  3. Older Codex versions disable tool features in certain deployment modes.

When integrating against models lacking tool calling, Codex cannot run file or shell operations, limiting it to plain chat workflows.

5.5 Session State and Stale Authentication

Symptoms: AuthError, session expired, conversation stops unexpectedly. Codex retains session state and cached authentication tokens. When configuration changes do not resolve errors, fully reset the login state:

codex logout
codex login

Custom provider keys and official OpenAI login tokens can conflict. Remove stale authentication files if credentials overlap between providers.

6. Security Best Practices and Operational Advice

6.1 Secure API Key and Conversation Data

Never hardcode API keys directly into config.toml. Use environment variables to inject secrets.

  • Restrict file permission: chmod 600 ~/.codex/config.toml.
  • Rotate keys immediately if configuration files are leaked to public repositories.
  • Do not share API keys across multiple team members; use separate scoped keys for each developer.

6.2 Backup and Version Control Workflow

Maintain copies of working config.toml profiles before modification. Version control configuration templates without embedded secrets.
Use codex resume to restore interrupted sessions, paired with model and CODEX_HOME switching to test different model providers against identical coding tasks. This is the most reliable way to compare model performance for specific engineering workloads.

6.3 Debugging Requests

When unexpected behavior appears, codex --debug prints full HTTP request and response payloads. Inspecting raw request bodies quickly reveals incorrect headers, path issues or malformed parameters.

Conclusion

Codex CLI becomes a flexible universal coding agent once users move past the default OpenAI configuration. The three core controls are model, model_provider, and provider blocks in config.toml. Engineers can route coding workloads to internal LLMs, local inference servers or third-party compatible endpoints without modifying the client source code.

The biggest source of integration issues is misunderstanding how model_provider resolves endpoints, combined with TOML syntax mistakes, protocol mismatches and stale authentication sessions. The debugging workflow recommended in this guide validates endpoints independently with curl before touching Codex configuration, isolating backend service problems from client configuration bugs.

For teams running multiple LLM backends, separating provider profiles with CODEX_HOME keeps environments clean. API gateways simplify credential management and traffic routing, and Treerouter delivers a unified layer to abstract model backend differences for Codex CLI and other agent clients.

Learn more:https://treerouter.com