Run DeepSeek V4 Flash in Codex CLI.

DeepSeek V4 Flash can act as the model behind the Codex CLI agent harness, but the connection needs a Responses-compatible route. The reliable path is LM Studio or another gateway that exposes /v1/responses; the official DeepSeek endpoint is not a direct drop-in for current Codex.

Quick answer: load a DeepSeek V4 Flash build in LM Studio, start its local server, copy the exact model identifier, then run codex --oss --local-provider lmstudio -m <model-id>. A successful text reply is only the first check; verify file and shell tool calls before trusting it with a long coding task.

What "DeepSeek on Codex" actually means.

Codex remains the agent

Codex still provides the terminal workflow, repository access, approvals, tools, and task loop. DeepSeek V4 Flash supplies the model responses.

It is not an OpenAI model

The standard Codex subscription and OpenAI model selector do not turn into DeepSeek. You are choosing a separate local or hosted inference provider.

The wire format matters

Current Codex expects the Responses API. DeepSeek officially documents Chat Completions and Anthropic-compatible endpoints, so a compatible gateway must bridge the boundary.

Tool use is the real test

A model can answer a prompt while still failing on tool schemas, reasoning replay, or the next turn after a command result.

Supported setup: LM Studio.

LM Studio documents a native Codex integration through its OpenAI-compatible POST /v1/responses endpoint. DeepSeek publishes V4 Flash weights and links to quantized builds compatible with local applications. The model is large, so choose a build your hardware can actually load.

Install a compatible DeepSeek V4 Flash build in LM Studio and load it with enough context for agent work.
Start the local API server on its default port.
List the served models and copy the identifier exactly; do not guess from the display name.
Launch Codex in OSS mode with LM Studio selected and pass that exact model identifier.
Run a read-only repository task, then one bounded edit and verification command before attempting a long autonomous job.
# Start LM Studio's local server
lms server start --port 1234

# Confirm the server and exact model identifier
curl http://localhost:1234/v1/models

# Run that model behind Codex CLI
codex --oss --local-provider lmstudio -m <exact-model-id>

Verify the Responses endpoint first.

If Codex reports a provider or streaming error, test the boundary without the agent loop. Replace the placeholder with the identifier returned by the model-list endpoint.

curl http://localhost:1234/v1/responses \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<exact-model-id>",
    "input": "Reply with exactly: READY",
    "stream": false
  }'

If this request fails, fix the server or model before debugging Codex. If it succeeds but Codex tools fail, the problem is tool compatibility rather than basic inference.

Using the official DeepSeek API.

The official model identifier is deepseek-v4-flash. DeepSeek currently documents it through OpenAI-style Chat Completions and an Anthropic-compatible interface. Current Codex custom providers use the Responses API, so pointing Codex directly at https://api.deepseek.com is not the safe setup.

To use the hosted DeepSeek API with Codex, place a Responses-compatible translation gateway in between. Configure Codex against the gateway, not directly against the Chat Completions URL.

model = "deepseek-v4-flash"
model_provider = "deepseek_responses_gateway"

[model_providers.deepseek_responses_gateway]
name = "DeepSeek through Responses gateway"
base_url = "http://127.0.0.1:PORT/v1"
env_key = "DEEPSEEK_API_KEY"
wire_api = "responses"
Do not copy this block unchanged: PORT and the gateway behavior depend on the translator you operate. Confirm that it preserves streaming, tool calls, tool results, and multi-turn reasoning data.

Fix the common failures.

Model not found

Use the exact ID returned by /v1/models. A catalog name, filename, and loaded model identifier may differ.

Connection refused

Start the server, confirm port 1234, and keep it bound to localhost unless remote access is deliberately secured.

Chat works, tools fail

Verify that the gateway supports custom tools through /v1/responses, not only plain text generation.

400 during a tool turn

DeepSeek reasoning modes can have stricter tool-choice and reasoning-history requirements. Test a simpler mode or fix the gateway's translation.

Context disappears

Increase the served context cautiously and send a compact working packet. The advertised maximum does not guarantee your local build can serve it efficiently.

Slow load or memory failure

Choose a smaller quantization, reduce context, or use hosted inference. Do not treat an out-of-memory server crash as a Codex bug.

Preserve the task before switching models.

A provider experiment should not make the coding state depend on one model finishing correctly. Before moving between an OpenAI model and DeepSeek V4 Flash, capture the current goal, changed files, git status, git diff, decisions, failed attempts, last passing check, and one next action.

ShardStitch prepares that state as a trust-labeled continuation packet. Codex, DeepSeek, or another supported target receives the same verified disk facts without treating the old model's narrative as source of truth.

Source trail.

This setup is grounded in current primary documentation:

FAQ.

Can Codex CLI run DeepSeek V4 Flash?

Yes. Codex CLI can use it when the model is served through LM Studio or another Responses-compatible endpoint. DeepSeek is the model provider; Codex remains the coding-agent harness.

Can I point Codex directly at the official DeepSeek API?

Not as a drop-in current Codex provider. DeepSeek documents Chat Completions and Anthropic-compatible interfaces, while Codex expects Responses. Use a compatible gateway or translator.

What command runs it through LM Studio?

Start the LM Studio server, retrieve the exact model ID, then run codex --oss --local-provider lmstudio -m <model-id>.

Why does the model answer but fail to edit files?

Plain inference is working, but the tool loop is not. Verify Responses tool support, tool-result replay, reasoning compatibility, context, and the exact gateway build.

Scope note: This guide covers Codex CLI as an open agent harness. It does not claim that DeepSeek V4 Flash is included in an OpenAI subscription, hosted by OpenAI, or selectable in the standard Codex cloud model list.