<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Load Balancing on Phạm Duy Tùng Machine Learning Blog</title>
		<link>https://www.phamduytung.com/tags/load-balancing/</link>
		<description>Recent content in Load Balancing on Phạm Duy Tùng Machine Learning Blog</description>
		<generator>Hugo</generator>
		<language>vi-VN</language>
		
		
		
			<copyright>Copyright © 2016-{year} Phạm Duy Tùng. All Rights Reserved.</copyright>
		
		
			<lastBuildDate>Tue, 18 Aug 2026 00:00:00 +0700</lastBuildDate>
		
			<atom:link href="https://www.phamduytung.com/tags/load-balancing/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>6,12 tỷ request LLM đã nói gì về Cache, Load Balancing và cách chúng ta đang phục vụ AI?</title>
				<link>https://www.phamduytung.com/blog/2026-08-18-llm-serving-cache-load-balancing/</link>
				<pubDate>Tue, 18 Aug 2026 00:00:00 +0700</pubDate>
				<guid>https://www.phamduytung.com/blog/2026-08-18-llm-serving-cache-load-balancing/</guid>
				<description>Phân tích paper A Year in LLM Serving từ hơn 6 tỷ request production: workload LLM không đứng yên, request count có thể đánh lừa, prefix cache mang tính thời gian rất mạnh và cache không thể được tối ưu tách rời khỏi load balancing.</description>
			</item>
	</channel>
</rss>
