Tested DeepSeek API compatibility: Tool Calling passed 3/3, but strict JSON Schema returned HTTP 400

I ran a small compatibility test against the official DeepSeek API using the deepseek-flash model ID. This wasn't a model-quality benchmark. I wanted to check how the API actually behaves with Chat Completions, SSE streaming, tool calling, and structured output. The main finding: Tool Calling and json_object worked in all three runs, while strict json_schema requests consistently returned HTTP 400. 1. Test setup Provider: DeepSeek Official API Model ID: deepseek-flash Endpoint: Official OpenAI-compatible API Basic connectivity checks: 1 run Deep capability tests: 3 runs Hidden retries: None Test location: Shanghai, China Network: Residential broadband / Wi-Fi This is a small sample from one network and location, not a long-term reliability or performance study. 2. Basic connectivity The single Basic run passed all checks: URL: PASS DNS: PASS TLS: PASS HTTP: PASS Authentication: PASS /models : PASS The /models endpoint returned a valid model list. The target model was also found during the Deep tests. 3. Chat Completions and SSE Streaming Both worked successfully across all three Deep runs. Capability Result Chat Completions PASS — 3/3 SSE Streaming PASS — 3/3 Streaming latency observations: Run TTFT Total latency 1 719 ms 985 ms 2 703 ms 1,015 ms 3 953 ms 1,031 ms Median 719 ms 1,015 ms These measurements are only point-in-time observations and shouldn't be interpreted as provider-wide performance figures. 4. Tool Calling: PASS — 3/3 All three runs successfully completed an actual tool-calling round trip. The test flow was: Send a request with a tool definition. Receive a valid tool call from the model. Validate the tool call and execute the local test tool. Send the tool result back to the model. Receive a successful second model response. Both API requests returned HTTP 200 in all three runs. This matters because accepting a request containing a tool definition isn't the same as successfully completing a tool-calling workflow. Observed result: SUPPORTED — 3/3 5. Structured Output: Different results for json_object and json_schema This was the most interesting part. json_object All three runs succeeded: HTTP 200 Request accepted Response parsed as valid JSON Returned value was a JSON object Result: SUPPORTED — 3/3 Strict json_schema All three runs returned: HTTP 400 Request rejected Classification: feature_rejected Result: Support status UNKNOWN The requests were consistently rejected, but I wouldn't conclude that strict JSON Schema is universally unsupported based on these three tests alone. The exact request format, parameters, endpoint variant, or model version could matter. Summary: Chat Completions PASS 3/3 SSE Streaming PASS 3/3 Tool Calling PASS 3/3 json_object PASS 3/3 strict json_schema HTTP 400 3/3 UNKNOWN support The key distinction: an API supporting json_object doesn't necessarily mean it supports strict json_schema with the same behavior. 6. Token usage reporting Provider-reported usage was available for Chat and Streaming in all three runs. Run Chat tokens Streaming tokens 1 85 71 2 79 60 3 71 57 These are prompt + completion token totals. Usage was not recorded separately for Tool Calling or Structured Output, so these numbers don't represent complete test costs. 7. Limitations A few important caveats: Only three Deep runs were performed. No hidden retries were used. The results don't establish long-term reliability or SLA. Tool Calling passing 3/3 doesn't imply a 100% success rate. The strict JSON Schema failures don't prove permanent or universal lack of support. Latency measurements aren't suitable for comparisons with other providers. Results may differ across accounts, regions, endpoints, or model versions. 8. Questions for other DeepSeek API users I'm particularly interested in whether others have seen the same structured-output behavior. Has anyone successfully used response_format: {"type": "json_schema"} with deepseek-flash on the official DeepSeek API? If yes, what request structure or parameters worked? Are you using json_object instead, or validating structured output at the application layer? I'd be interested in comparing actual request payloads and HTTP responses to understand whether this is a model-level limitation, an endpoint-specific behavior, or a request-format compatibility issue.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论