<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Prefix Caching on AI Tech Blog</title>
    <link>https://jesamkim.github.io/ai-tech-blog/tags/prefix-caching/</link>
    <description>Recent content in Prefix Caching on AI Tech Blog</description>
    <generator>Hugo -- 0.147.6</generator>
    <language>ko</language>
    <lastBuildDate>Fri, 11 Sep 2026 21:44:07 +0900</lastBuildDate>
    <atom:link href="https://jesamkim.github.io/ai-tech-blog/tags/prefix-caching/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>캐시를 켰는데도 LLM이 느린 이유: Prefix-aware Routing의 원리</title>
      <link>https://jesamkim.github.io/ai-tech-blog/posts/prefix-aware-routing/</link>
      <pubDate>Fri, 11 Sep 2026 21:44:07 +0900</pubDate>
      <guid>https://jesamkim.github.io/ai-tech-blog/posts/prefix-aware-routing/</guid>
      <description>LLM의 prefix cache를 여러 서버에서 효율적으로 재사용하려면 요청 라우팅도 함께 설계해야 합니다. SageMaker의 prefix-aware routing을 사례로 캐시 적중률과 부하 분산, 첫 토큰 지연의 관계를 살펴봅니다.</description>
    </item>
  </channel>
</rss>
