Curated coverage, model releases, technical papers, and engineering updates tagged with #Deepseek.
A deep dive into the DeepSeek-V4.1 Flash technical report: CED, CSA2, HSI, Single-Pass mHC, Engram and FP4 KV Cache — how the KV cache was compressed to just 890 bytes per token.
DeepSeek’s 11/11 result showed why advanced agent benchmarks need to check both the outcome and the attack path: our audit confirmed six planned exploits and found five unexpected routes.
PyTorch implementations of modern open-source LLM architectures (Llama, Qwen, DeepSeek, Gemma, GPT-OSS, Kimi, and more) — written from scratch for readability and learning, based on Sebastian Raschka's LLM Architecture Gallery. - anuj0456/OpenArc...
got deepseek V4.1 flash running locally on a 16GB m1 mac mini original FP4/FP8 weights, ssd streaming + custom mlx runner 108s ttft and about 23s/token (not to be confused with tok/s)
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Introducing the smallest model in our new architecture family, with native visual understanding. Designed for greater capability, faster inference, and higher throughput.
A new report released Thursday by Anthropic alleges persistent distillation attacks by China-based AI companies, which have escalated in recent months as competition in the space has intensified.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Debut closes in on DeepSeek's trillion-parameter flagship.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is …
The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model. I …
You've reached the end of the dispatches.