senn-techsenn-tech
← IT News
Week 2026-W332026-08-16

IT News KW33/2026: GPT-5.6 Gets Cheaper, DeepSeek V4-Flash Surpasses Claude Fable 5, Apple Screen Sharing Wide Open

AIOpenAIDeepSeekQwenSecurityApple

The week of 10–16 August 2026 revolved around two themes: prices are falling, and vulnerabilities are being exploited immediately. OpenAI cut GPT-5.6 Luna by 80% and turned speed into a separate product. DeepSeek V4-Flash gained 25 points on Terminal-Bench and now leads Claude Fable 5. Apple had to patch a critical authentication bypass in macOS Screen Sharing, and Alibaba positioned Qwen3.8-Max as an open frontier alternative.

Deep Dive: OpenAI GPT-5.6: Price Collapse and Speed-as-a-Service

Luna Down 80%, Sol Faster for a Surcharge

OpenAI cut prices for the cheaper GPT-5.6 models: Luna now costs $0.20 input and $1.20 output per million tokens (−80%). Terra is at $2 / $12 (−20%), Sol stays at $5 / $30. A Fast Mode for Sol was added, offering up to 2.5× speed at double the price. the decoder

The message: speed is decoupled from the model and becomes a separate product. Those who can wait pay less. For SMBs, this means planning assumptions age quarterly, and long-term contracts based on current prices become expensive quickly.

Assessment: Lower prices are good, but only for those who remain exchangeable. Anyone locked into a single model misses the cheaper alternatives.


Deep Dive: DeepSeek V4-Flash: 99% Cheaper Than Claude Fable 5

Silent Retraining, Big Benchmark Jump

DeepSeek retrained V4-Flash without fanfare. Same architecture (284 billion parameters, ~13 billion active, 1 million token context), new post-training. On Terminal Bench 2.1 the model jumped from 56.9 to 82.7 points: ahead of Claude Fable 5 (80.5). OfficeChai

The price is $0.14 input / $0.28 output per million tokens. Claude Fable 5 charges $10 / $50 for the same amount. The performance gap to the top model GPT-5.6 Sol is only 3.1 points, while the price difference is a factor of 100.

Terminal-Bench 2.1, Score und Listenpreis (Input / Output je 1 Mio. Token)GPT-5.6 Sol85.8$5 / $30DeepSeek V4-Flash82.7$0.14 / $0.28Claude Fable 580.5$10 / $50Claude Sonnet 574.5$3 / $156090
DeepSeek V4-Flash liegt 3,1 Punkte hinter GPT-5.6 Sol und vor Claude Fable 5, zu rund einem Prozent von dessen Preis. Quelle: Terminal-Bench 2.1, Anbieter-Preislisten (Juli 2026).

Assessment: For agentic coding tasks, V4-Flash is now the first choice when data privacy is not a concern. Those with sensitive codebases should use open weights in self-hosting.


Deep Dive: Apple macOS CVE-2026-65400: Screen Sharing Without a Password

Authentication Bypass Allows Root Access

Apple patched CVE-2026-65400 on 6 August. The flaw is in screensharingd and allows an attacker with network access to authenticate to the VNC service on port 5900, without valid credentials. Result: root access via remote desktop. NCSC-NL confirmed active exploitation; in every known case, a Monero miner was installed. Tanium

Apple raised the severity from 7.1 to CVSS 9.8. CISA added it to the KEV catalog on 18 August.

Assessment: Screen Sharing is active in many companies for support and is rarely checked for reachability. Patch immediately, disable Screen Sharing where it is not needed, and block port 5900 externally.


Deep Dive: Qwen3.8-Max: 2.4 Trillion Parameters With Open-Weight Promise

Alibaba Returns to Open-Sourcing Its Top Model

Alibaba released Qwen3.8-Max on 3 August: a MoE model with 2.4 trillion parameters, ~95 billion active, multimodal for text, image, video, and documents, with 1 million token context. The price is $2 input / $6 output per million tokens. MarkTechPost

In the Arena leaderboard for frontend code, Qwen3.8-Max ranks 4th with 1,668 points, 37 points behind Claude Opus 5. More important than the ranking is the promise to release the weights: both Qwen3.8-Max and a smaller Qwen3.8-27B. If that happens, a 2.4-trillion-parameter model would become self-hostable.

Assessment: The announcement is vaguer than the weights themselves. Until checkpoints are available, Qwen3.8-Max remains an API option. Organizations with privacy requirements should wait for self-hosting.


Innovation & Open Source

  • FreeToken: New Apache 2.0 edge-native MoE inference engine that runs 290B+ frontier models on a gaming GPU plus CPU/RAM. Competition for KTransformers.
  • vLLM 0.27.1 prefix-token details: --enable-prompt-tokens-details makes the cache share per prompt transparent.

Digest: Other Key News

Security

AI & Enterprise


From the Blog


Compiled on 16 August 2026. Sources: the decoder, OpenRouter, OfficeChai, Tanium, MarkTechPost, Bloomberg, Cyber Recaps.