#2 Everyone's Trying to Cage Their Own AI Right Now

#newsletter #ai

Building a RAG pipeline from scratch

I built a RAG pipeline from scratch, end to end, and wanted to share how it works under the hood.

The pipeline has six steps: chunk the source document, embed each chunk into a vector, store those vectors, retrieve the closest matches for a question, generate a grounded answer, and evaluate the whole thing with an LLM as judge.

For my demo, I used a Wikipedia page on the FIFA World Cup, split into 89 chunks (1,000 characters each, 200-character overlap), embedded everything locally with a compact sentence-transformer model, and stored the vectors in ChromaDB, which is a free, local vector database. Retrieval and generation ran through Groq's free-tier LLaMA 3.3 endpoint.

The part I'm most proud of is the evaluation stage. Instead of reaching for a library, I built a simple LLM-as-judge setup that scores each answer on three axes: faithfulness (is it grounded in the retrieved context or hallucinated), answer relevancy (does it actually address the question), and context precision (are the retrieved chunks useful or just noise). It's a clean, minimal way to actually know whether your RAG system works.

Building a RAG Pipeline from Scratch: Chunking, Embeddings, Vector Storage, and LLM-as-Judge…A hands-on guide to chunking, embedding, and evaluating with LLM-as-JudgeMedium

A slide-deck template library for your coding agent

If you ever build slide decks with a coding agent, this repo is worth bookmarking. It's a library of 30 free, MIT-licensed HTML slide templates, each with its own visual identity. It contains everything from a soft editorial serif look, to an 8-bit pixel-art aesthetic, to a full Windows 95 chrome theme. Definitely worth taking a look.

The idea is simple: instead of asking your agent to design a deck from a blank page, point it at this library. An accompanying instructions file tells the agent how to read the template index, match it to your brief, clone the right one, and adapt the content so you end up with something that actually looks designed, instead of generic.

zarazhangrui/beautiful-html-templatesA library of HTML slide templates designed so any coding agent can pick the right one and produce a beautiful deck on the user's behalf, automatically.GitHub

Multimodal prompting is quietly becoming the new normal

There's a video making the rounds about multimodal prompting, and the argument is worth summarizing: prompt engineering isn't dying, it's evolving into something a lot richer than a well-worded sentence.

The idea is that developers are moving past text-only instructions and steering AI coding agents with a mix of inputs: screenshots, mood boards, UI sketches, even hand-drawn mind maps. Instead of optimizing one static line of text, you guide the agent through the whole loop of planning, prototyping, and debugging using much richer context.

Two things stood out. First, human expertise doesn't get replaced here, it becomes more important, since domain knowledge is what actually lets you judge whether the agent's output is any good. Second, the back-and-forth itself, brainstorming and visualizing ideas side by side with the agent, tends to surface the "unknown unknowns" you'd never have thought to type into a prompt in the first place.

My new way of working with AI Agents.
My new way of working with AI Agents. · Elvis Saravia

Fable 5 is back. So is the backlash.

Anthropic's most powerful model has had a rough few weeks. Fable 5 launched on June 9, got pulled three days later over Commerce Department export-control concerns, and came back on July 1 once those controls were lifted.

The comeback came with a catch. To get cleared, Anthropic built a stricter safety classifier targeting a jailbreak technique an Amazon researcher had reported. Anthropic says the new classifier blocks that specific technique in over 99% of attempts. The tradeoff: it's also flagging a lot of ordinary coding and debugging requests, quietly rerouting them to the less capable Opus 4.8 instead.

Independent benchmarking group BridgeMind reran its coding tests on the relaunched model and reported some steep drops. Debugging scores fell from 86.2 to 25.9, refactoring from 73.6 to 38.4. Developers on Reddit and X described the same pattern, complaining that everyday tasks, including plenty with no real safety angle, were getting bounced to the weaker model.

Anthropic maintains the underlying model hasn't changed, and the numbers seem to back that up: on the smaller share of tasks that don't get flagged, Fable 5 performs about the way it did before the ban. The real issue is the routing itself. It's a good reminder that "the model got weaker" and "the guardrails got stricter" can look identical from the outside and in this case, it's the latter.

On top of the classifier changes, access is capped too: Fable 5 is limited to 50% of a user's weekly usage through July 12, after which it shifts to a pay-per-credit system for everyone.


From GPU basics to writing your own inference kernels

If a career in AI infrastructure is on your radar, this GitHub repo is close to a full syllabus. It's a tiered curriculum, from GPU fundamentals all the way up to the optimizations used at frontier AI labs, organized into three levels of depth across several tracks: hardware architecture, matrix multiplication and tensor cores, memory-bound kernels like FlashAttention and PagedAttention, compiler approaches like Triton and CUTLASS, production inference systems such as vLLM and SGLang, and even a newer track on AI agents that write their own GPU kernels.

It's maintained by Wafer, a company building AI-assisted tools for profiling and optimizing GPU kernels. Whether you're starting from zero or already comfortable with CUDA, it's a solid map of where to go next.

wafer-ai/gpu-perf-engineering-resourcesA curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.GitHub

Stop asking the judge for a score. Ask it yes-or-no questions instead.

Grading LLM outputs with a single score from another LLM is common practice, but it has a real weakness: one number tells you almost nothing about what actually went wrong.

A new paper called BINEVAL proposes a fix. Instead of asking a model to rate an output on a 1–5 scale, it breaks each evaluation criterion down into small yes/no questions, then aggregates the answers into a score. Instead of "rate this summary's factual consistency," you'd ask a series of concrete questions whether every named entity is accurate, whether any numbers were misstated, whether the summary invents information that isn't in the source.

Across three established benchmarks, the decomposed approach matched or beat existing evaluators, and it did especially well at catching subtle factual errors. The researchers also used the same question-level feedback to automatically improve prompts, both for the evaluator itself and for the model doing the generating, by having a stronger model flag the specific questions where a weaker one disagreed.

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-ImprovementEvaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that…arXiv.org

China might be drafting its own version of the AI export ban

Two weeks after the US briefly pulled Anthropic's Fable 5 and Mythos 5 models over export-control concerns, China appears to be sketching out a version of the same playbook.

Reports say China's Ministry of Commerce has spent the past month in talks with Alibaba, ByteDance, and startup Z.ai about restricting overseas access to their most advanced AI models, including some not yet released. The proposals reportedly go beyond a simple export ban. They'd cover both closed and open-weight models, meaning systems like Alibaba's Qwen, ByteDance's Doubao, and Z.ai's GLM-5.2 could all be affected. One version floated a tiered system: basic open-source tools under light filing requirements, more advanced technology under security review, and the most capable frontier models restricted to domestic use only.

Nothing has been decided, and the officials involved haven't commented publicly. But it lands amid rising friction between the two countries' AI sectors, Alibaba reportedly banned employees from using Claude Code starting this week, after a developer found code in the tool designed to detect whether a user was connected to a Chinese AI lab, something Anthropic said was aimed at combating model distillation.

If both governments end up restricting their own frontier models this way, the open, largely borderless flow of AI capability that's defined the last couple of years might start closing in on itself, from both sides at once.


The tiny model making a big case for running AI at home

While the big labs argue over export controls, there's a quieter trend worth paying attention to: models getting small enough to just run on your own hardware.

Liquid AI's LFM2.5-VL-450M is a good example. It's a vision-language model with only 450 million parameters. Tiny next to the 7B-plus models that usually handle image tasks and it runs entirely on a phone or a laptop. Despite the size, it handles visual question answering, object detection with bounding boxes, and structured data extraction straight into JSON.

Part of what makes this possible is architectural. Liquid AI skips the standard Transformer design in favor of a non-Transformer architecture based on dynamical systems, which turns out to be notably faster and lighter on memory, especially for long context or on-device use.

Given how much of this issue is about who controls access to the big models, a genuinely capable model you can run locally with no API, no rate limit, no guardrail that might change overnight starts to look less like a novelty and more like a hedge.

LFM2.5-VL-450M : This Tiny AI is Breaking the Internet 🤯 (Runs Locally!)
LFM2.5-VL-450M : This Tiny AI is Breaking the Internet 🤯 (Runs Locally!) · Codedigipt