Manufactured Sources Behind AI Recommendations — Trellner Research

Ask an AI assistant which CRM to buy and it usually grounds the answer — retrieves web pages first, then summarizes them. Trellner Research wanted to see what actually fills that evidence base. They asked two Perplexity models for the best product in 380 categories, and logged every one of the 7,534 web pages the models pulled in as support. The results are uncomfortable for anyone who trusts AI search: ...

September 2, 2026 · 2 min

Instinct vs Grok Bot vs ChatGPT vs Hermes: Which AI Agent Can You Trust? — Peter Yang

Peter Yang runs four personal AI agents — Instinct, Grok Bot, ChatGPT/Codex, and Hermes — that can read his email, open his documents, use his logins, and make purchases. ~24 minutes of live demos and a trust audit of each. The lineup Instinct — invite-only agent that lives inside iMessage/WhatsApp as a single thread; founder reportedly raising at a $2.5B valuation; “the closest thing to texting a trusted friend who can just do stuff for you” Grok Bot — a team of named bots on a persistent 24/7 cloud computer that can hand information to each other ChatGPT + Codex — still where 90%+ of his real work happens: ChatGPT runs tasks in the cloud, Codex works on local files Hermes — the open-source agent running on his Mac mini, talked to via Telegram; “gives me the most control” Instinct — magical but opaque Simple: no threads or bots to manage; connecting Google Workspace was a one-tap OAuth link Resourceful: found his Google AI Ultra free trial would auto-renew at $100/month, browsed the gift terms to confirm he could cancel renewal and keep the free year, then cancelled it — but the cancellation required him to hand over his 2FA code and Google password into a password vault form Proactive: uses cron jobs and scheduled tasks behind the scenes — it emailed his golf instructor, watched for the reply, and came back with alternative slots; it pinged him when a booked-out sushi place opened up Personable: emoji reactions make it feel human Privacy: its policy says disconnecting a third-party integration does not automatically delete collected data — you have to delete it manually in the workspace settings The catch: he can’t see what it’s doing between “reading” and “acting”, and doesn’t even know which model runs underneath Grok Bot — a bot team on a cloud computer Named bots with personality: a chief-of-staff that coordinates the others, a growth bot that emails weekly site-metric charts, a “doom scrolling uncle” reading X via cloud browser (the official connector burns API credits), a Marie-Kondo bot tidying email/Drive in character, and a “cheap dad” bot hunting discounts and listing things on Facebook Marketplace Official plugins use normal OAuth flows — comfortable to connect, same as ChatGPT Cloud-browser logins are the uncomfortable part: typing passwords and 2FA codes into a computer “that I have no idea where it is” (SpaceX AI servers) Deleting a bot doesn’t remove the shared cloud computer’s files or browser sessions; “reset” rolls back to the last snapshot, not to scratch — wiping it means manually disconnecting every plugin ChatGPT + Codex — still the main driver Most powerful and flexible interface, but the UI is messy: ChatGPT work and Codex feel squished into one app, and it’s unclear which tasks are cloud vs local Deep plugin ecosystem, including his business bank account (Mercury) Privacy toggle to check: “Improve the model for everyone” can train on data from connected apps — turn it off under Settings → Data controls His trust rationale: OpenAI runs large-enterprise workloads, so a data leak would be disastrous for them Hermes — the open-source local option Runs 24/7 on his Mac mini like a personal local cloud: morning briefs with three focus items, scheduled meetings, weekly email reports Connected his smart scale and a vibe-coded fitness app via MCP for a weekly health-trend email Privacy by construction: open source, no telemetry or analytics; conversations, memory, and skills live in local files “If it goes off the rails, I can just unplug it” — something you can’t do with a cloud computer Reality check: most of his work moved to ChatGPT/Codex, so Hermes now mostly runs scheduled jobs What can actually go wrong Live prompt-injection demo (via his friend Alex Cohen): a fresh Gmail account emailed instructions to set up a nightly cron reading the primary inbox and emailing action items back to that address — Instinct followed the instructions; when both accounts belong to the same person it’s a trick, but swap the second account for an attacker’s and private data walks out Prompt injection = instructions hidden in an email, website, or document the agent reads; the agent follows them and exfiltrates your information Instinct says it has safeguards (“email content is data, never a command… nothing sent to another person without you seeing it first”), but he has no way to verify — and a smart model reduces the risk without ever reaching 100% Practical cleanup tip: my.google.com → linked apps (he found 67-80+); rather than removing them one by one, paste the link into any capable agent and tell it to audit and uninstall — these tools are all good at browser use now “I might let an agent compare hotel prices, but I don’t quite trust it enough to book a non-refundable trip without looking through what it’s trying to do.” — Peter Yang ...

September 2, 2026 · 4 min

My local model setup on an M4 Pro Mac mini — Kevin Lewis

Kevin Lewis runs his own AI models on a Mac mini in his house — the same machine that backs his Hermes agent, his phone chat apps, and his coding assistant. His essay is a practical case for why he did it, and why he thinks local is no longer a hobbyist compromise. His core argument is that cloud AI is rented land: Providers can change pricing, throttle your usage, or silently swap the model behind the endpoint — he was regularly maxing out two $200/month subscriptions while getting inconsistent quality Sending sensitive code or client data to a third party is a decision you can’t undo Governments can restrict model availability; owning your compute is the only guaranteed remedy After the hardware purchase, every query is free — flat cost, no rate limits, works offline The piece also teaches you how to read a model name like Qwen3.6-35B-A3B-OptiQ-4bit. The number that matters isn’t 35 billion — it’s the ~3 billion parameters actually activated per token. Mixture-of-experts models spread weights across many “experts” but only wake a few at a time, which is why a 35B-class model fits in 20GB of memory. Compression to 4-bit precision costs only a couple of benchmark points versus the full-quality baseline. ...

September 2, 2026 · 2 min

EFF to Courts: Don't Rewrite Copyright over AI Hype — EFF

Every new creative technology triggers a copyright panic. In the 1980s the VCR was called “the Boston strangler” of the film industry. Before that, the player piano was going to destroy music composition, and cameras were going to kill portrait painting. None of it happened — photography ended up birthing entire new art forms like photojournalism. The EFF’s latest DeepLinks essay argues the AI copyright wave is the same story, and courts should treat it the same way the Supreme Court treated the VCR: with skepticism about hype. ...

September 1, 2026 · 2 min

How accurate have Ed Zitron's AI skeptic predictions been? — Dan Luu

Dan Luu — the systems engineer behind some of the most-cited teardowns on the internet — decided to check the record of Ed Zitron, the AI skeptic tech media quote most often. So he went through Zitron’s predictions one by one and graded them. The verdict: wrong on roughly everything, and wrong on the reasoning, not just the outcome. “Peak AI” called repeatedly from Feb 2024 through 2025 — models kept improving each time Meta, Google, and Microsoft called “dying” in late 2024 — all grew revenue and profit at double-digit rates through 2025 and the first half of 2026 OpenAI’s growth “stalling” — it exceeded its own revenue forecasts Gemini hitting 500M users “so unrealistic that someone should be fired” — it passed 750M Cursor “going to die” — acquired for $60B; CoreWeave “can’t survive six months” — more than doubled its IPO price Luu’s point isn’t that the numbers are merely wrong. It’s that they’re deployed as rhetorical cover: a spreadsheet that double-counts a month, third-party traffic data that contradicts the company’s own reports, and a style that floods you with so much confident nonsense that refuting it costs more than producing it — a tactic known as a gish gallop. ...

September 1, 2026 · 2 min

How to Turn Your AI Into a World-Class Designer — Anshu Chimala

Anshu Chimala — 12 years leading software engineering and design at Apple, now posting AI-design demos on X — argues the reason AI churns out “generic slop” for you and “magic” for him isn’t the model. LLMs are next-token predictors trained to make the safe, average choice every time, which is design-by-committee. The fix is a reimagined Double Diamond process — Discover → Define → Deliver — tuned for a team of AI agents. ...

September 1, 2026 · 2 min

11 Tiny Coding Agent Fixes With a Stupid Amount of Payoff — Cole Medin

Cole Medin runs through 11 small tweaks that make any coding agent — Claude Code, Codex, whatever — noticeably more reliable, without scrapping your workflow. The through-line: agents are prediction machines, not deterministic programs, so reliability comes from shrinking their decision space and moving guarantees into deterministic mechanisms. ~17 minutes. Rules and context 1. Write for the agent, not the human. Humans interpret high-level docs in context; agents make assumptions. Be blunt — file paths, numbers, exact commands. 2. Your instruction files rot. “Rule drift”: 1 in 4 repos have stale AI rules referencing deleted files or replaced databases. Audit them against the codebase. 5. Less context is more. Too many rules hurts as models improve. Keep global rules under ~200–300 lines; drop generic advice and move the rest to task-specific context files. Conversation hygiene 3. /compact loses ~90% of detail. Compacting a bloated conversation breeds hallucination. Give smaller work chunks, or write your own handoff doc and start fresh. 7. Don’t escalate mid-task. A bigger model can’t rescue a tainted conversation — mistakes compound within a session. Write a handoff doc and burn it. 10. Over-revision degrades quality. 85% of the time an earlier iteration was better. The model “fixes” things just to appease you. Determinism over frameworks 4. Load-bearing rules → hooks. Rules are probabilistic (the agent will “forget” to run tests); hooks are deterministic — fire on an event and route failures back. 6. Subagents eat your rate limit. Parallel fanouts cost more than you think — 39% of his weekly usage came from 4+ parallel sessions. 8. You don’t need coordinators. Team-lead frameworks and agent mailboxes are unreliable. A plain delegator agent gets most of the scale with far more reliability. Validation 9. Never let the writer approve the work. The writer carries its own bias. Review in a fresh conversation with a handoff doc — no assumptions carried over. 11. Validation is a system, not a step. Plan the full harness — test conventions, tools, edge cases — before writing any code, not as an afterthought. “Your number one job when you’re planning any work with your coding agent is to reduce the number of assumptions that it’s making.” — Cole Medin ...

September 1, 2026 · 2 min

AI News - 2026-09-01

Tuesday’s digest ran on the largest collection this pipeline has seen — 1,761 items, with arXiv alone at 1,487 — and it had a clear lead: Anthropic’s first-party postmortem of the summer’s agent-incident thread, with escape-attempt classifiers, a multi-week RL pause, and new rules for third-party evaluators. The rest of the day split between agent security and self-host inference: covert indirect prompt injection and a skill-injection threat model landed on arXiv, DeepSeek shipped the V4-Flash family’s first open vision weights, and a budget-aware pipeline squeezes a 70B onto a single GPU at ~33GB. Simon Willison published a companion reference site for ChatGPT Work, and Z.ai’s H1 numbers put a concrete figure on open-model API economics. ...

September 1, 2026 · 6 min

Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance

Pete Johnson — MongoDB’s field CTO of AI and a 30-year database veteran — makes the case across ~95 minutes on The Cognitive Revolution that the interesting frontier in AI has moved back into database territory. His thesis, stated once: agent performance, and especially cost-adjusted agent performance, depends on retrieval — what you choose to put in front of the model, in what order. The thesis Not the model, not the context window, not the prompt — retrieval is what decides whether an agent is good Everything else in the conversation (database history, the Voyage acquisition, vector search) is downstream of that one claim A history of constraints 1970: SQL is born (E.F. Codd, IBM) — storage was the scarce resource, so normalization (store nothing twice) was the right design 2007: MongoDB’s first commit, the year the iPhone ships — after 47 years of Moore’s Law, time became scarce, so denormalize: one JSON document, one disk read instead of three “The problem has faded, but the solution persists” — Johnson agrees, and flags the education system’s “thou shalt always normalize” bias MongoDB’s AI path — it started with keyword search 2020: customers stood up their own Lucene servers for keyword search → MongoDB shipped Atlas Search (lexical, managed) A vector is just an array of floats → to a document DB, that’s just another attribute, so vector search was cheap to add Three levers: pre-filter (metadata) + lexical + vector = hybrid search in one query (rank fusion / score fusion, one API call) 2025: the Voyage acquisition — and the conversation pivots from “database features” to “embeddings actually matter” Embeddings are not commoditized “Most people think embedding models are commoditized — that is not true” Hugging Face’s Rtech benchmark: up to a 14% quality gap vs. the default picks 14% is the difference between a hallucination and a correct answer A reranker on top adds another 5–10% ($re-rank, one-call) Anthropic — no embedding model of its own — recommends Voyage Three Voyage features that remove plumbing Contextualized chunking: send the chunk plus its surrounding context, get one vector back — better retrieval at smaller chunks, inverting the normal tradeoff Matryoshka reasoning: dimensions nest like Russian dolls — embed at 1024, lop off the last 512 to test, no re-embedding your corpus Shared embedding spaces: four sizes of one model share an embedding space; a free open-weight “nano” can run queries locally to kill token cost in dev The memory problem, compressed 2022: query → context window → answer. 2023: the knowledge cutoff + proprietary data → RAG. 2025: tools/MCP + looping → the memory problem Early answer: short-term memory = cram the session; long-term = cram the last three days Two failures: token maxing (Uber burned its entire 2026 budget in 13 weeks) and lost-in-the-middle (the first and last ~7K tokens are what matter; the middle muddies the answer) The fix is selection, not stuffing Stop asking “how do I cram a million tokens in” — ask “how do I choose the right 200K for this loop” Taxonomic memory: a hundred company-specific terms exist, but only five are relevant to this loop — pick those five, re-pick next loop Two responsibilities now: query with a token budget, and write the answer back so the system curates and stores it Write, change, recall, forget Memories have a half-life — recent matters more — and forgetting is the hardest part Nathan’s own memory system (monthly logs → yearly summaries → entity wiki) hits both pain points: the DRY violation and the model keeping a dead project open for months Guidance: a good embedder + reranker makes the forget step workable; graph structure for the top 2–6 levels, vector search in the leaf; don’t run multiple LLM passes to shrink the corpus — that’s just more tokens Memory done well: ElevenLabs’ micro-agents, one per customer Build vs. buy, three camps Camp one: “I bought one tool, I’m done.” Camp two: POC purgatory — usually the wrong problem. Camp three: optimizing sophisticated memory Problem selection: top 10–15 problems, which have good data, which already have metrics — else you can’t tell if AI helped “Bad data quality and bad security posture don’t get solved by AI — they get amplified” Lines of code is a terrible metric; idea-to-production is the one that matters The world outside the US Seven countries, ~100 customers this year — and the two most sophisticated were in Mexico City and São Paulo, both assuming US competitors were ahead Nearly every country has a hyperscaler data center now — the geographic barriers that kept US companies ahead have eroded “We’ve been building databases for 60 years. We’ve been building agents for about 18 months… there’s no LAMP stack for agents yet — no React and Angular, no established right answer an enterprise can confidently buy.” ...

September 1, 2026 · 4 min

Stop Shipping AI Nobody Can Verify — Hamel Husain

Hamel Husain (Parlance Labs, evals course co-author with Shreya Shankar) on Vanishing Gradients with Hugo Bowne-Anderson — ~76 minutes on why verification should drive AI product design, and how evals changed once agents arrived. The whole conversation orbits his recent post, “It’s hard to eval is actually a product smell.” ...

September 1, 2026 · 4 min

Stop Picking Embedding Models Off The MTEB Leaderboard — Radu (Vespa)

Radu, a search engineer at Vespa, on how to actually choose and tune embedding models for search and RAG — hosted by Hamel Husain. ~22 minutes. Why not just take the leaderboard Most people pick embeddings one of two ways: the top of a benchmark leaderboard, or the model from whatever provider they already use — leaving a lot on the table Cost, latency, robustness to future model changes, and your actual use case aren’t reflected in a single ranking MTEB is a starting point, not an answer: find the subtask closest to your use case, restrict by model size, and watch for small models that punch above their weight (this talk covers dense embedders — one vector per chunk) What MTEB doesn’t tell you Quantization depends on where you run: on CPU, int8 keeps most of the precision at a fraction of the cost and runs much faster; on GPU, an int8 model runs at its native rate and is actually slower — run FP16 instead, basically the same result for significantly less Vector precision: storing vectors as bf16 instead of float32 is a no-brainer (no measurable loss); bit vectors (packed to int8) do cost quality — but in hybrid search the gap shrinks and they become viable Matryoshka dimensions: some models are trained so the earliest dimensions matter most, so you can just cut the vector — 2048 → 1024 dims cost nothing, 512 a little, 256 more; and 2048-dim bit vectors run a quarter the size of bf16 at 512 dims Query latency comes from two places: how fast the embedder turns the query into a vector, and the distance math — normalized vectors let you use dot product instead of cosine, bit vectors unlock Hamming distance (fastest, and CPU-optimized in modern search engines) Constraints MTEB never covers: multilingual needs, long-context support Tune on your own data with Vespa Embed Open-source fine-tuning tool (UI over sentence-transformers, models from Hugging Face): feed pairs (query → document) or triplets when you have negatives; it auto-splits a validation set MNR (multiple negatives ranking) treats other documents in the batch as negatives; symmetric and JSD variants check whether a batch-picked document is actually a negative Start with defaults; if you have labeled negatives, try triplets — his e-commerce results were similar either way, but real hard negatives (from search logs: positives rank on top, negatives below) are worth mining No clean labels? Use an LLM as a judge: have it rank query→document pairs on a 0/1/2 scale, give it examples (the most important part), and run micro-batches of 8-10 docs so it doesn’t assume listwise context Fine-tune embeddings far more readily than LLMs — it’s a constrained problem: cheap, quick, and consistently a big NDCG jump The cheap-embedder escape hatch: re-rankers Two-phase search: an intentionally “great but not greatest” cheap embedder does first-phase ranking over millions of docs, then a more expensive re-ranker (float vectors, cross-encoders, late interaction) scores only the top N Store cheap bit vectors (with HNSW) in memory for the first phase and keep float vectors on disk for the re-rank — memory is expensive, disks aren’t Late-interaction re-ranking (MaxSim) is literally: per-token dot products, keep the max per patch, sum across the document Do’s and don’ts from real experiments Don’t train on title + description as a proxy for queries — proxies lie; people’s real queries look nothing like product copy. Train on real queries and relevance judgments NDCG has a blind spot: change your relevance function and newly-surfaced documents score zero until they’re rated — if you don’t re-rate everything in the metric, your NDCG looks like it never improves Pick the metric by the problem: e-commerce ranking → NDCG/ERR (top results matter); RAG → precision matters more, because junk retrieved = hallucination, and recall is capped by context size Don’t over-optimize: with ~2M documents most of these knobs don’t matter yet — start with defaults that lose little, then measure “The embedding is a much more constrained problem — and it’s a lot easier to fine-tune. I always find a really good benefit from doing it. And it’s not that complicated.” — Hamel Husain ...

August 31, 2026 · 4 min

Agent memory as a file format — Cal Paterson

Cal Paterson’s thesis is simple: AI agents should start with memories, not a blank slate — but almost every agent-memory system on the market gets it wrong. His fix is a file format, not a framework. He groups today’s memory systems into three failing camps: Harness-locked memory that mines your chat history — mostly remembers things about you instead of the world Complicated pipelines — one prominent system needs a vector database, a graph database, and its own LLM just to decide what’s worth remembering “High Modernist” memory — distilled facts and graphs that strip information from its context until it’s senseless His alternative, “memoryfields,” is a folder of Markdown pages plus an optional search index that finds pages by meaning (semantic search) rather than keywords. Agents write memories directly in Markdown — their favorite format — instead of feeding text through chunking and summarization machinery. ...

August 31, 2026 · 2 min

Breaking Claude Code Opus 5 Auto Mode — Johann Rehberger

Claude Code now runs in Auto Mode by default. Instead of asking a human before each command, a safety classifier decides what’s allowed. Anthropic commissioned a third-party evaluation that reported a 0.00% prompt-injection success rate for Opus 5 in Auto Mode — and security researcher Johann Rehberger wanted to see if that held up against a targeted attack. Prompt injection means slipping hidden instructions into content an AI agent reads; here, a website that asked Claude to summarize a page. The attack didn’t rely on “ignore your instructions” tricks — it made the malicious path look like the natural one: ...

August 31, 2026 · 2 min

Understanding ChatGPT Work — Simon Willison

Simon Willison spent weeks reverse-engineering ChatGPT Work, the agentic product OpenAI launched in July. His conclusion: it is really two products — Work Cloud (in the browser and mobile apps) and Work Local (the renamed Codex desktop app) — and the cloud version is the one worth understanding. It is also $20/month and up only. What separates Work from plain ChatGPT Chat: Code execution with full internet access — it can clone GitHub repositories, install dependencies, and talk to any API. Chat’s sandbox blocks that; even Claude’s container allows only a short list of sites A full headless Chrome browser — it loads pages, fills out forms, and takes screenshots. When a site needs a login, it can hand over to you for passwords and two-factor codes without those secrets ever passing through the model A persistent filesystem shared across sessions — files from one chat stay available in the next. Willison already has 171 scratch folders ChatGPT Sites — it can build and deploy real websites on Cloudflare Workers, databases included, from a single prompt Sub-agents — parallel model sessions working on one project — plus scheduled prompts that check things for you on a timer The demo that sells it: one prompt asked Work to find every “pelican in her piety” in London, turn the results into a JSON file, and build a website about them. It did all of it, end to end, from one instruction. ...

August 31, 2026 · 2 min

A Claude Cowork System That Does a Week of PM Work in a Day — Daniel Bloom (How I AI)

Daniel Bloom, a PM at Melio (fintech), shows Claire Vo his Claude Cowork system on How I AI — a personal agent harness that manages his week: “I’m able to do in a day what used to take me a week.” The two rules of a powerful system It can rewrite its own core files — the system keeps improving itself It connects to as much of your ecosystem as possible — Notion, calendar, Slack, Gmail, Granola meeting transcripts The tool matters less than these two properties — Cowork works for him, but Codex or ChatGPT Work could do the same The architecture Notion as a read-only brain: three columns (Top of Mind / This Week / Inbox) — Cowork built the board itself when it got tired of his messy Google Doc, and manages it on his behalf Context files: a CLAUDE.md-style context file for every work area, goal, and colleague; he spent the first weeks “contextualizing ruthlessly” — feeding links, decks, and endless voice notes (Whisper) Weekly prep (Sunday): a recurring task composed of skills — pulls his whole ecosystem, recommends the week’s focus, triages meetings into serious-prep / quick-reminder / nothing Daily brief (morning): walks through yesterday’s meetings via Granola transcripts with one-liners and action items, then asks what to expand or draft The killer feature: proactive context The daily brief scans recent Slack/email/notes for things it doesn’t understand — new files, milestones, goals, internal terms — and asks him to define them, then saves them to context. Internal vocabulary like “settlement cap” never makes it into training data; this is how the system learns the company’s real language. Claire’s verdict: “really sharp, something we haven’t seen on the podcast.” ...

August 31, 2026 · 3 min

AI News - 2026-08-31

Monday’s digest skews agent security: three fresh papers — long-context prompt injection, quantization-triggered backdoors, and out-of-band policy enforcement at a trusted tool boundary — plus the first concrete signal that agent pricing is moving from tokens toward outcomes. The tooling side delivered too: OpenClaw 2.0 is the platform’s largest release yet, and Simon Willison published the most complete hands-on map of ChatGPT Work so far — useful before you spend on a $20/mo sub. Five arXiv papers and the outcome-based pricing shift round it out. ...

August 31, 2026 · 4 min

The One Skill That Survives The AI Shift — Ofer Mendelevitch

Ofer Mendelevitch (Vectara; author of Hands-On RAG for Production and, with Jay Alammar, Hands-On Large Language Models) interviewed by Angelina on TwoSetAI — 70 minutes on production RAG and surviving the AI shift. RAG isn’t dead — it just got a loop RAG = retrieval + augmented generation, and retrieval stays essential even with agents; the “RAG is dead” claims every two months are mostly people marketing something new Classic RAG is one-shot: query → top chunks → prompt → answer. Agents add a loop: the LLM plans, calls tools (often the RAG pipeline itself) repeatedly, and synthesizes “Talk to my PDF” demos are not production: millions of documents in every format change the problem completely The pipeline, from ingest to answer Ingest: extract text → chunk → embed → vector store (store the text and page markers too, not just vectors — citations need to point at exact pages) Hybrid search (semantic + BM25/TF-IDF) for what semantic search misses: numbers, product codes, exact strings — “if the source says 90% and your output says 85%, vector search won’t catch it” Query side: top-5/10 results → optional reranking → prompt → grounded answer Tables, images, and the red button problem Tables are first-class citizens: chunked tables lose their column names (common in medical journals) — store whole tables, retrieve them whole Images: store and return as images, don’t flatten to a description Video: transcription alone loses meaning — “in order to avoid catastrophe, never press this button” means nothing without the visual. Today’s fix: VLM descriptions of short clips correlated with the transcript. Dedicated video embedding models exist but aren’t production-ready yet When to bother with knowledge graphs Multi-hop questions (“what else did the director of Inception direct?”) defeat semantic search But graphs are expensive to build and maintain — worth it only for high-stakes use cases where a significant share of queries actually need the relationships Eval: the hard part is the data, not the metric Retrieval eval (did you fetch the right chunks?) needs query→gold-chunk datasets that are brutal to build — and documents keep changing Generation eval compares against curated golden responses Reference-free eval (Jimmy Lin’s Waterloo lab + Vectara): LLM-as-judge scorers like UMBRELLA (0–3 chunk relevance) validated against human correlation — no gold labels required Build, buy, or rent Build with LangChain/LlamaIndex only if it’s your business and you have the team; complexity compounds (multimodal, graphs, maintenance) RAG/agent-as-a-service (e.g. Vectara) outsources the upkeep; vertical tools are fine when they cover your use case — but watch missing features and data-residency constraints DevRel as a growth engine Two jobs: teach developers how to use the product, and carry feedback back to the company PLG over expensive sales teams: self-serve product, events, hackathons, real blog posts — “make it your own voice, don’t produce AI slop” Measure directionally, don’t over-engineer attribution: five customers means it’s not working, ten thousand means it is, a thousand is unknowable — same problem founders face reading PMF The skill that survives Engineers and data scientists become directors, not actors: agents write the code; the remaining critical challenge is deciding what to build and steering where agents are weak (architecture, non-obvious trade-offs) To the high schoolers who feared they made “the most incredibly stupidest mistake” by majoring in CS: graduate with the capability of today’s mid/senior engineer — use college to learn how to wield the AI tools Fundamentals still matter; hiring will have to change — “write Fibonacci in five lines of Python is worthless” — expect AI-augmented interviews His real worry is societal, not technical: how governments and finance distribute the gains “We’re going to end up in engineering and data science being directors as opposed to actors. The coding agents will write the code.” ...

August 30, 2026 · 3 min

AI News - 2026-08-30

Sunday was a quiet one — four items, no arXiv feed (weekend skip), no fresh cross-platform launch. The day’s biggest story is day-2 fallout from yesterday’s OpenAI–Cursor split: Cursor co-founder Michael Truell says OpenAI was only ~5% of traffic and that Cursor trusted it to stay “neutral,” with Musk shrugging it off. The real artifact on the stack is a detailed break of Claude Code’s Auto Mode — 60–80% injection success against Anthropic’s commissioned 0.00% claim, attack chain fully written up. Around it: the largest music-industry copyright suit yet against a lab, and a first-party change to Claude Code’s weekly usage limits. ...

August 30, 2026 · 3 min

No AI Fridays — Carson Gross

Carson Gross, the creator of HTMX, has mandated “No AI Fridays” at his company: one day a week where AI assistants stay off and the team writes code by hand, reads the documentation, and thinks things through. It’s a one-page manifesto that hit the top of Hacker News and kicked off a long debate. The case against always-on AI: The costs are documented: studies link heavy LLM use to “cognitive debt” (accepting output without internalizing it), weaker critical thinking, less engagement with work, and slower skill formation. Blind spots: when a model makes the decisions, you stop noticing the trade-offs — a day off is a chance to check the AI is still steering you where you want to go. Missed automation: defaulting to AI skips “good old automation” — sometimes the right tool is a small script you write yourself, not a token-hungry model. The piece is deliberately deadpan — the FAQ reads like an intervention script (“A couple of shots of Claude or a pint of Codex is always best when you try to quit, right?”) — but the argument underneath is serious and fits a growing literature. ...

August 30, 2026 · 2 min

How Non-Coders Are Vibe Coding $100K+ Businesses With AI — Amol Jain (Replit)

Peter Yang interviews Amol Jain, Replit’s head of engineering, on the businesses people actually build with vibe coding — and what that means for SaaS. ~36 minutes. Real businesses built by non-coders Pep AI — Cedric, a 22-year-old Oregon student who had never coded: a peptide/GLP-1 tracking app; $60K in his first month, tens of thousands of active users AI proficiency platform — John, a repeat founder quoted $100K+ by an agency: built end-to-end in 3 days on Replit, $180K+ revenue within 2 months, purely on product intuition and distribution skill TryNearby — FaZe Apex’s Yusuf: matching local creators with local businesses for word-of-mouth; YC Summer 26 batch, $100K+ ARR What the successful ones share A unique advantage turned into a product: expertise, judgment/taste, community, or distribution — “Replit takes a person’s advantage and turns it into a business” Not massive TAMs — small thriving businesses serving a few hundred customers are the norm “Users are people and inherently lazy. If you offer them value, they will pay for it.” Generic use cases (meal planners, exercise trackers) are commoditized — agents will swallow them whole Vibe-coded to production The gap between a raw vibe-coded app and something production-grade: Replit’s flow is publish + domain, then Security Center (an agent builds a threat model of your app and runs a deep scan — secret scanning, malicious packages, PII leakage; previously a 1-2 week, contracted security-team job), then Stripe monetization (the platform sets up a Stripe sandbox, plans, and prices for you — you just claim the account), then an SEO scan tuned for both search and answer engines Culture: employees build internal tools constantly — a RevOps person with no coding background built the sales demo tool in 2 days (replacing a six-figure SaaS product), a data scientist (not a coder) built their in-house BI analytics “A 10x person is building a tool that the entire team or company can use” — leverage for everyone What survives the “cost of code goes to zero” test The question: if code is free, what do you pay for? Answer: trust, data, infrastructure, atoms (physical things), labor, and network Salesforce survives as the system of record (headless, with everyone building on top); Workday survives because payroll is regulated; DoorDash survives because humans deliver food; social networks survive because you can’t buy the network Existing SaaS is going headless (APIs, CLIs, MCPs) — and agents can now operate browsers anyway App layer vs model labs Models are commoditizing; being model-agnostic is a feature (latest-and-greatest, cost — “token-maxing is coming to an end”, and resilience when a single provider goes down) The mental model: a frontier model does the hard part, an orchestrator delegates to sub-agents on cheaper models Enterprises care about ROI, not leaderboards — “the R part has become very relevant” “If the cost of code goes to zero, what do you pay for? It’s this unique expertise… it’s the unique judgment or taste. It’s the security. It’s the distribution. Users pay for that.” ...

August 30, 2026 · 3 min