
Jev's Boring Use Cases Are the Ones That Actually Work
Jev is a lightweight model that runs far faster and cheaper than larger LLMs, delivering competitive accuracy for tasks such as reranking, j...
Subscribe to Lê Đức Minh's Newsletter
Get new articles and engineering posts delivered to your inbox.

Jev is a lightweight model that runs far faster and cheaper than larger LLMs, delivering competitive accuracy for tasks such as reranking, j...

Reproducing Laya’s RLCD shows that its “RL” term is just a noise‑smoothed estimate of the cross‑entropy gradient, which makes inference over...

Laya's README reports a raw ECE of 0.466 on the typed-decisions benchmark before temperature scaling.

I went into this expecting the RL to be the interesting part. Laya's whole pitch is that a reinforcement-learning term turns noisy logits in...

Jev, a decision‑only AI model released on September 15 2026, was quickly adopted by developers: pi‑warden, built within 48 hours, blocked 42...

RLCD is a proprietary training method from TypeSafe AI that teaches Jev to produce calibrated probability distributions and confidence score...

Jev is a “System One Model” that returns typed, calibrated probabilities for predefined question types in a single forward pass, enabling fa...

Connect a vision-language model to a live RTSP surveillance feed, ask it to generate real-time incident summaries, and watch your GPU metric...

Give an autonomous coding agent a hundred-thousand-token context window, point it at a GitHub repository, and ask it to reproduce an ML base...

Dynamic reasoning budgets route simple tasks to a zero‑thought path and reserve capped reasoning for complex queries, cutting token use and ...

Learn how a deterministic harness prunes context, uses a ledger and transactional tool calls to keep LLM agents reliable over many turns.

A few weeks ago, my dad and I spent an afternoon at the kitchen table trying to figure out what an "AI agent" actually is.

State‑of‑the‑art LLMs, even Vietnamese‑specialized ones, score below 50 % on the VIVID benchmark of 1,636 authentic Vietnamese idioms and fa...

Benchmark reveals Llama-3.1-8B mislabels Vietnamese slang as angry, losing 20 F1 points, while three other LLMs correctly interpret social m...

Discover why Vietnamese tokenizers cost only 1.05‑2.14× English tokens, not 4.5×, and how updating tokenizers can cut LLM expenses by a thir...

TokPress compresses tiny JSON log lines by tokenizing with OpenAI's o200k_base tokenizer, then applying LZ77 and rANS for smaller files.

TokPress compresses tiny JSON log lines by tokenizing with OpenAI's o200k_base tokenizer, then applying LZ77 and rANS for smaller files.

Field notes from a weekend spent teaching an entropy coder to speak LLM. LLMs spent billions of dollars learning the best subword dictionary...

In July I wrote Prompt, Context, Harness, Loop, which split an agent into four parts and argued that the harness — the thing that owns the l...

Part 2 of 2 on prompt caching. Part 1 covered the economics. Part 1 established the prize: caching cut an 80-turn agent session from $54.08 ...

Part 1 of 2 on prompt caching. I got curious about a number I had never actually checked: what does a long session with a coding agent reall...

Field notes, part 1 of 3. Data plumbing from an AI engineer's desk.* My job title says AI. A meaningful share of my week is data. Not the gl...