Curated coverage, model releases, technical papers, and engineering updates tagged with #LLM.
Muse is getting a lot of attention — including mine — because it’s both groundbreaking technically (each user gets their own entire persistent Linux VM running in Meta’s cloud) and …
NarrateAI delivers production-ready LLM quality assurance on Amazon Bedrock. This post details five techniques—adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, composite evaluation, and data accuracy verification—that reach about 99% numerical accuracy while streaming...
The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder. We can do amazing things with them, but unlocking …
Which LLM is worth it: Artificial Analysis Intelligence Index against blended API price, with the value frontier highlighted. Refreshed daily.
Nori LLM: the fastest large language model on the market. 1,000,000+ tokens per second. Optimized for humans and robot crawlers alike.
Analysis of Inception's Mercury 2.5 and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
I'm hosting an evening event with Jesse Vincent in San Francisco on Wednesday 14th October for people who are building weird and interesting things with and on top of coding …
Yesterday was Grok 4.7 (pelicans) and MiMo v2.6 Flash/Pro (more pelicans). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It’s …
Hey, you know it's like super obvious if you're using AI to write your scripts for TikTok and YouTube, right? [...] It's not just the general AI-isms of "it's not …
Hi HN, I built ai·rete·rag because I kept seeing teams put an LLM in charge of decisions that need to be auditable (lending, fraud, clinical triage), then bolt on "guardrails" after the fact.It runs the two in series instead:1. A pure-Python Rete engine evaluates YAML r...
LLM agents can write a lot of code for you very fast, but they also take away opportunities to learn and grow.
Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling “System One models” (I’m with Maggie Appleton, I think “decision models” …
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
Agent evals and guardrails in one request. Built on Jev, Kev and Laya. - openlayer-ai/jevals
It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, …
The un-bannable, checksum-verified mirror for sovereign AI models.
The Rust compiler remains unperturbed by the antics of the LLM
Gemini finally caught up on Felony Bench! The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was …
Being a computer scientist who refuses to find anything about LLMs interesting right now is a bit like being a geneticist who refuses to find anything interesting about the recently …