In July 2026, Chinese AI inference has reached $0.07 per million tokens — against $1.25 for comparable Western models — and Goldman Sachs has begun recommending Chinese models directly to Wall Street clients, according to AI in China's tracker. The eighteen-to-one pricing differential has shifted the conversation from geopolitical hesitation to active procurement evaluation at major financial institutions.
The trigger is structural: DeepSeek V4 Pro runs at $0.87 per million tokens off-peak versus Anthropic's $50 for equivalent Fable 5 volume, while Alibaba's Qwen 3.8 preview is priced at a tenth of standard Alibaba rates. Neither DeepSeek nor Alibaba has published independent third-party benchmarks for their latest releases, leaving buyers to weigh cost certainty against capability uncertainty.
The Goldman recommendation is the most concrete signal yet that export-control policy has not prevented cost arbitrage from reshaping enterprise AI sourcing. If sustained, it creates a compliance and reputational question for U.S. financial regulators that is distinct from — and potentially harder to contain than — direct government procurement bans.
As of July 2026, a Chinese AI model costs $0.07 per million tokens while a Western equivalent is priced at $1.25 — an 18-to-1 spread that has prompted Goldman Sachs to actively steer Wall Street clients toward Chinese models. The gap reflects both deliberate pricing strategy and the compounding efficiency gains Chinese labs have extracted from constrained compute.
The dynamic pricing DeepSeek introduced alongside V4's GA launch — $0.87 per million tokens off-peak for V4 Pro — sits at the higher end of the Chinese market. Smaller inference providers running Qwen and DeepSeek derivatives are already offering rates well below that floor, using open-weight models that require no licensing fees.
Roughly 41% of all Hugging Face Hub downloads over the trailing year came from China-origin models. The pricing compression is structural: as long as open-weight releases continue at the current pace, Western closed-model providers face sustained margin pressure with no obvious floor in sight.
DeepSeek's V4 model moved from preview to general availability on Sunday July 20, completing a three-month release cycle that began with its initial drop in April. The lab paired the GA launch with a dynamic pricing structure: during off-peak hours, V4 Pro is available at $0.87 per million tokens — versus roughly $50 for the same volume from Anthropic's Fable 5, a 57× gap.
No formal benchmark suite accompanied the release. DeepSeek framed the pricing as the lead signal, consistent with its established strategy of making capability accessible at drastically reduced cost. The V4 Pro architecture builds on the MoE (mixture-of-experts) design that powered V3, with the 'Pro' variant understood to be the full-scale deployment version.
The launch landed alongside Alibaba's Qwen 3.8 preview hours later, compressing what would normally be separate news cycles into a single competitive day for the global open-model market.