Thursday 13 August 2026 | Join Free | Upgrade

Hi there, this is your daily ☕️ AIpresso.

In today's newsletter:

🤖 AI ran first autonomous hack on Taiwan

🤖 Grok 4.6 matches GPT-5.6

🤖 Nvidia's new Nemotron model bets on speed

👁️ Liquid AI vision model runs on edge

🐋 DeepSeek upgrades its V4 Pro model

Plus: 🎁 6 other news you might like, 🛠️ 5 engineering & practice picks, 🧰 6 tools, and 📚 5 papers.

See how AI delivers fast, secure, personalized experiences

Join Fin and Plaid on August 13th to see how AI can resolve issues directly in customer conversations. This includes everything from reducing bank-linking friction, enabling proactive support, and helping customers complete real financial tasks.

🤖 AI ran first autonomous hack on Taiwan LINK
  • Suspected Chinese hackers ran the first publicly known near-autonomous AI cyberattack on a government, using open-source models to breach Taiwanese state infrastructure and extract over 2,500 personnel records, according to research from Israeli cyber firm Dream.
  • The multi-agent system, built on the Hermes and OpenClaw frameworks, ran "Learning Cycles" that autonomously searched vulnerability databases, GitHub, and security publications, then expanded to IT supply chain vendors, a nuclear safety agency, and 7+ energy companies in parallel.
  • Attackers bypassed safety guardrails by framing the operation as authorized penetration testing, though Dream stressed the system still required substantial human work, task-specific tuning, agent coordination, and decision-logic adjustment, rather than "just" running a model.
🤖 Grok 4.6 matches GPT-5.6 LINK
  • SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing only Claude Opus 5 (63) and Fable 5 (62), a five-point gain over Grok 4.5.
  • The model shines on agentic work, ranking second on GDPval-AA v2 with an Elo of 1,753 and finishing complex tasks in ~53 steps versus Claude Opus 5's ~103.
  • Pricing holds at $2/$6 per million tokens, over 60% below Claude Opus 5 and GPT-5.6 Sol, and it's live via API, Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare, with doubled quota the first week.
🤖 Nvidia's new Nemotron model bets on speed LINK
  • NVIDIA shipped Nemotron 3.5 Lightning on August 11, a 30B open-weight model distilled from its 550B Nemotron 3 Ultra and built for the execution layer of agent loops rather than general-purpose reasoning.
  • Independent tests clock it near 670 tokens/second-roughly double Gemini 3.5 Flash-Lite-with strong agent scores like 824 Elo on GDPval-AA v2, priced on OpenRouter at $0.05/$0.20 per million tokens under an unrestricted-commercial OpenMDW-1.1 license.
  • Its Intelligence Index of 24 trails Qwen 3.6-35B-A3B (32) and Gemini 3.5 Flash-Lite (37) on MMLU-Pro, GPQA Diamond, and SWE-bench Verified, though its 1M-token window retrieves at just 52.0 accuracy, behind smaller-context rivals.
👁️ Liquid AI vision model runs on edge LINK
  • Liquid AI released LFM2.5-VL-3B, a 3B vision-language model that runs fully on-device, decoding 228 tokens/s on an M5 Max, 20 tokens/s on a Galaxy S26 Ultra, and fitting in about 3 GB of memory.
  • Built on a SigLIP2 400M NaFlex encoder and the LFM2.5-2.6B backbone, it was pre-trained on ~34T tokens with 4x more vision data, and leads its size class on screen/UI understanding, grounding, and document reading.
  • The model hits ~11K output tokens/s at high concurrency on an H100, roughly 2x larger 4B-class models, and ships day-one on llama.cpp, MLX, vLLM, SGLang, and ONNX via Hugging Face.
🐋 DeepSeek upgrades its V4 Pro model LINK
  • DeepSeek has quietly shipped V4 Pro 0813, its latest reasoning model, available API-only through OpenRouter with no formal announcement page from DeepSeek itself.
  • The model exposes three distinct reasoning levels, low, medium, and high, which produced noticeably different outputs, an unusual behavior compared to other models tested.
  • Open weights haven't been confirmed but seem likely given prior V4-Pro and V4-Flash-0731 releases, though benchmark numbers only surfaced secondhand via WeChat, a deleted Reddit post, and a Hacker News ASCII table.

The first way to trade directly inside Claude and ChatGPT

Superintelligence used to be locked inside billion-dollar quant firms whose algorithms quietly took advantage of everyone else. 

Co-Invest puts it right in your chat window. Analyze markets, manage risk, and execute trades, all inside Claude and ChatGPT. 

The institutions built the game, Co-Invest gives you a way to beat them.

🛠️ Engineering & Practice

> One attention head carries knight forks in a chess transformer, and here's a new toolkit that found it.: A new open-source toolkit for chess AI interpretability pinpointed that a single attention head handles knight-fork tactics, proving specific skills live in specific components.
> Role boundary plasticity: prompt injection gauntlet reveals 12 of 16 frontier models will Wire A stranger your $500: Frontier AI models can't reliably tell trusted user commands from text sneaked into tool results, so 12 of 16 fired fake $500 refunds-and size or price doesn't predict safety.
> When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results: Injecting false concepts into an AI model shows it only notices the tampering after it has already given a wrong answer, never before.
> Introducing the conceptual reasoning index: The Conceptual Reasoning Index scores how well AI models argue about unverifiable safety questions, so labs can track a skill critical to reducing AI risk.
> 11X cheaper than ChatGPT: Tiny 150M model just proved AI doesn't need to "think out loud" to be smart: Pathway's tiny AI reasons inside its memory instead of writing out its steps, matching bigger rivals at roughly eleven times lower cost per task.
 

Other news & articles you might like

  • Unsloth desktop brings local AI training to Mac, Windows and Linux LINK
  • Anthropic brings claude cowork to its chrome extension, adding skills and plugins to the browser LINK
  • OpenAI’s ChatGPT/Codex desktop app is now on Linux LINK
  • Okta targets AI agent token costs with MCP scoping LINK
  • How One startup Is bringing Wall Street–style contracts for AI tokens LINK
  • Google debuts SL2T, an AI model that’s designed to understand sign language LINK

📚 Trending papers & reports

> Multi-task AI training gets a smarter way to balance competing goals, using the natural matrix shape of models like Transformers instead of flattening them, consistently boosting both training speed and final task performance. LINK
> Chatbot persuasion disclosures only weaken a chatbot's influence when it reveals its persuasive intent, cutting attitude shifts from ~13 points to ~6 points, while merely labeling it "AI" changes nothing. LINK
> Crowd-powered AI serving lets everyday users' spare computing power handle overflow demand alongside company servers, cutting dedicated infrastructure needs while improving response completion and P99 latency as user numbers grow. LINK
> Nanophotonic design software turns a light-absorption target directly into an accurate physical structure blueprint, cutting reconstruction error and boosting shape accuracy versus standard methods, speeding development of optical materials and sensors. LINK
> Quantum image generation can produce sharper pictures using fewer quantum bits by generating each pixel individually from its coordinates, beating prior quantum methods on image quality while needing less hardware. LINK
 

🧰 Tools & repos

Soloop: an agentic system pairing solo founders with AI CEO, CTO, and CMO roles to plan, build, and sell without hiring a team. LINK
Portfolio Lab: tests AI-generated investment strategies on unseen data and live markets before letting you deploy vetted ones to your brokerage account. LINK
SecondBrain Note by GenSpark: an AI agent that researches topics across the web and compiles unbiased, well-organized Sparkpages, saving you time versus sifting through SEO-driven search results. LINK
AI Group Call: lets you state a goal and join a live voice call with six AI participants who discuss it, pause when you speak, and save transcripts and summaries. LINK
Dormice: a self-hosted, E2B-compatible sandbox environment for AI agents that runs persistently on a single machine with zero idle costs. LINK
Chat Agent by Trigger.dev: a template for building conversational AI agents in TypeScript, handling long-running tasks with automatic retries, queuing, and scaling. LINK

You can check the previous tools here, or add your tool here

🎓 Want to master the AI tools we cover every day?

Our AI Academy has 330+ step-by-step tutorials on ChatGPT, Claude, Perplexity, and every tool that matters. No fluff — just practical workflows you can use at work. Try it free for 7 days.

💬 How did you find today's edition?

We read every reply — just reply to this email and let us know how we can improve!

★★★★★  Nailed it
★★★  Average
  Fail

Not subscribed to ☕️ AIpresso yet? Subscribe for free