How semantic caching reduces LLM API calls

Bill Doerrfeld | May 5, 2025

My latest for The New Stack explores semantic caching, an emerging strategy to optimize agentic AI.


Semantic caching is like typical caching, but for AI. It could eliminate a lot of redundant API calls to LLMs, reducing costs and improving performance.


My latest for The New Stack explores semantic caching — what it is, how it works, and what the benefits are. According to the sources, semantic caching is poised to become more of a standard practice for optimizing how applications behave with AI, reducing latency and lowering the bar as costs increase.


Featured image credit: Donald Wu


Read: What Is Semantic Caching?

Other Blog Posts

By Bill Doerrfeld • October 5, 2026
My latest for LeadDev interviews a Meta engineer who built an agentic system to optimize their capacity planning, leading to real power use savings.
By Bill Doerrfeld • September 30, 2026
I set out expecting to write an article about how every enterprise was jumping on trendy alternative or "neo" clouds. That ended up not being the case.
By Bill Doerrfeld • September 28, 2026
Join the virtual AI Pen-Testing Conference to explore how AI is changing penetration testing, red teaming, and the security testing of AI systems.
By Bill Doerrfeld • September 28, 2026
Counting down to Nordic APIs Summit 2026 in Stockholm, featuring 60+ speakers exploring agentic AI, APIs, MCP, and security.
By Bill Doerrfeld • September 28, 2026
I covered a Yale and GitHub report on LeadDev that explores developer attitudes on climate change and curbing IT waste.
By Bill Doerrfeld • September 15, 2026
Does your agent really need access to every tool call under the sun?
By Bill Doerrfeld • September 8, 2026
I'll be writing a weekly series for The New Stack this fall covering Kubernetes, cloud native, and everything happening on the road to KubeCon in November.
By Bill Doerrfeld • September 1, 2026
I'm teaming up with Dan Barahona and AI Security University to put together the MCP Security Conference, a free half-day virtual event dedicated to MCP security .
By Bill Doerrfeld • August 27, 2026
My latest DirectorPlus edition interviews Shopify's head of engineering, Farhan Thawar, on their unique internal engineering culture.
By Bill Doerrfeld • August 24, 2026
Neoclouds promise cheaper GPUs and AI-optimized infrastructure, but can they really challenge hyperscalers? My feature on InfoWorld explores their advantages, risks, and likely role in AI compute.