Pi vs. the official DeepSeek Harness on the same local model (Qwen3.8-27B)
Pi outperformed the official DeepSeek Harness on a local Qwen3.8-27B model by passing 6 of 8 benchmark runs while using 14 percent less total execution time.
By Dillip Chowdary β’ Oct 10, 2026 β’ Source: github.com
According to benchmark results published in github.com's report, the Pi coding agent led the official DeepSeek Harness (DSH) in a 16-run coding-only evaluation on a local Qwen3.8-27B model. Pi completed 6 of 8 externally checked runs for a 75 percent pass rate, compared to 4 of 8 runs (50 percent) for DSH. Across the sample, Pi achieved a mean external score of 0.825 while DSH scored 0.680. Pi recorded two paired wins, zero losses, and six ties against DSH across four test fixtures.
This comparison covers performance metrics, resource consumption, token efficiency, and harness configurations for Pi version 0.73.1 and DeepSeek Harness version 0.1.1-rc.2. The evaluation targets software engineers, local LLM practitioners, and AI tooling developers evaluating local agent performance. Both harnesses ran on a single Apple M4 Max MacBook Pro with 128 GB of unified memory, utilizing an identical local oMLX server hosting the Qwen3.8-27B-MLX-oQ8e-mtp model.
The test: Pi vs the official DeepSeek Harness
The benchmark evaluated both coding agents using four specific tasks across two trials each, producing eight paired cells. The workload consisted of make-ci-green, add-feature, taskflow, and webcore. Both harnesses interacted with the same oMLX endpoint running Qwen3.8-27B-MLX-oQ8e-mtp under macOS 26.5.2 and Node.js 25.6.1. The model was configured with a 98,304-token context window, a 32,768-token output limit, a temperature of 0.6, and top-p of 0.95.
To ensure consistency, both harnesses ran under a common 1,200-second per-cell wall-clock execution limit. DeepSeek Harness was configured via version 0.1.1-rc.2 using the npm package @deepseek-ai/dsh@0.1.1-rc.2 and the @deepseek-ai/dsh-llm-pi-ai adapter. The benchmark disabled repository instructions, context files, skills, web search, session title generation, and model fan-out features in DSH to match Pi's single-attempt evaluation constraints and local proxy boundary.
How Pi and the official DeepSeek Harness each did
On the task level, both harnesses hit maximum performance ceilings on two fixtures and a floor on one. Both Pi and DSH scored 2 of 2 passes with 1.00 mean scores on make-ci-green and add-feature. Neither harness passed webcore (0 of 2 passes, 0.30 score), as neither agent edited any files during those runs. The separation between the two frameworks stemmed entirely from the taskflow fixture, where Pi achieved 2 of 2 passes while DSH failed both attempts, scoring 0.28 and 0.56.

In terms of execution speed and stability, Pi completed its runs faster while experiencing fewer timeouts. Pi logged 3 total timeouts compared to 4 for DSH. On the four paired runs where both agents passed, Pi executed faster by margins between 144.3 and 291.8 seconds. Pi required a total wall time of 6,455.5 seconds across all runs, whereas DSH required 7,356.5 seconds, representing 14.0 percent more wall-clock time.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
Pi vs the official DeepSeek Harness, side by side
A side-by-side analysis demonstrates differences in latency, token output, and generation throughput between the two agent architectures. DSH consumed 1,130,071 completed-response input tokens compared to Pi's 1,009,881 input tokens, marking an 11.9 percent increase for DSH. DSH also produced 30,520 completed-response output tokens against Pi's 26,587 output tokens, generating 14.8 percent more completed output text while completing fewer tasks successfully.
Pi demonstrated superior generation speeds and shorter response initiation times. Pi recorded an aggregate generation throughput of 12.74 tokens per second compared to 11.38 tokens per second for DSH, giving Pi a 10.7 percent advantage in generation speed. Pi's mean time to first token was 31.40 seconds, whereas DSH averaged 38.87 seconds, representing a 23.8 percent higher initial latency for DSH.
| Metric | Pi 0.73.1 | DSH 0.1.1-rc.2 | DSH vs Pi |
|---|---|---|---|
| External passes | 6/8 (75%) | 4/8 (50%) | -2 passes |
| Mean external score | 0.825 | 0.680 | -0.145 |
| Paired wins / losses / ties | 2 / 0 / 6 | 0 / 2 / 6 | Pi +2 wins |
| Timeouts | 3 | 4 | +1 |
| Total wall time | 6,455.5 s | 7,356.5 s | +14.0% |
| Median wall time | 906.5 s | 1,055.2 s | +16.4% |
| Endpoint requests | 86 | 80 | -7.0% |
| Completed responses | 83 | 76 | -8.4% |
| Completed-response input tokens | 1,009,881 | 1,130,071 | +11.9% |
| Completed-response output tokens | 26,587 | 30,520 | +14.8% |
| Median output tokens per run | 3,012 | 2,981.5 | -1.0% |
| Aggregate prompt tokens/s | 387.47 | 382.55 | -1.3% |
| Aggregate generation tokens/s | 12.74 | 11.38 | -10.7% |
| Mean time to first token | 31.40 s | 38.87 s | +23.8% |
See the Pi vs the official DeepSeek Harness output
Detailed task-level output reveals variance in tool usage and execution behavior under identical conditions. On Trial 1 of make-ci-green, Pi achieved a 1.00 score in 229.7 seconds using 6 requests and 1,253 output tokens across 6 changed files, while DSH achieved a 1.00 score in 374.0 seconds using 7 requests and 2,499 output tokens across 6 changed files. On Trial 1 of taskflow, Pi scored 1.00 in 1,194.6 seconds with 17 requests and 5,693 output tokens across 7 changed files, whereas DSH reached the 1,200-second cap with a 0.28 score after 12 requests and 3,586 output tokens without changing any files.
On Trial 2 of add-feature, Pi passed with a 1.00 score in 618.5 seconds using 14 requests and 5,804 output tokens across 2 changed files. DSH also passed with a 1.00 score but required 910.3 seconds, 10 requests, and 7,253 output tokens across 2 changed files. Throughout all requests, DSH exposed 19 model-facing tool schemas per request to the underlying model, whereas Pi operated with its 4 stock built-in tool schemas.
The verdict on Pi vs the official DeepSeek Harness
The benchmark results confirm that Pi led the official DeepSeek Harness on this specific 16-run local test sample. Pi demonstrated higher pass rates, lower total latency, lower token overhead, and higher generation throughput when connected to a local Qwen3.8-27B model running on an Apple M4 Max system. Pi's performance advantage was concentrated in the taskflow fixture, which served as the sole separating test case between the two agent frameworks.
However, the report notes that these outcomes do not establish that Pi is universally superior to DeepSeek Harness. Because the sample size was limited to four test fixtures across two trials, only two paired cells produced non-tied results. Additionally, DeepSeek Harness version 0.1.1-rc.2 is a developer preview release undergoing active development. The small sample size and fixture distribution prevent broad statistical rankings across general software engineering workloads.
Developer Action Items
- β Verify the claim on the official Apple / GitHub / macOS page (or HN AI Agents), not from this recap alone.
- β Name the surface that moved β API, policy, model, hardware, or commercial terms β before you Slack the thread.
- β Assign one owner a day to read the primary material and decide: this-sprint, this-quarter, or noise.
- β Do not change production on day-one coverage. Watch the vendor changelog and one independent write-up first.
Pi vs the official DeepSeek Harness FAQ
What was the pass rate comparison between Pi and the official DeepSeek Harness?
Pi passed 6 out of 8 externally checked benchmark runs for a 75 percent pass rate, while DeepSeek Harness passed 4 out of 8 runs for a 50 percent pass rate.
Which local model and hardware were used for the Pi vs DeepSeek Harness benchmark?
Both harnesses were tested on a local Qwen3.8-27B-MLX-oQ8e-mtp model running via an oMLX server on an Apple M4 Max MacBook Pro with 128 GB of unified memory.
How did execution speed and token usage compare between the two harnesses?
DeepSeek Harness required 14.0 percent more wall-clock time, produced 14.8 percent more output tokens, and had 10.7 percent lower generation throughput than Pi.
Sources
Author
Dillip Chowdary
Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.
Related on Tech Bytes
Advertisement