<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>AirLLM on Phạm Duy Tùng Machine Learning Blog</title>
		<link>https://www.phamduytung.com/tags/airllm/</link>
		<description>Recent content in AirLLM on Phạm Duy Tùng Machine Learning Blog</description>
		<generator>Hugo</generator>
		<language>vi-VN</language>
		
		
		
			<copyright>Copyright © 2016-{year} Phạm Duy Tùng. All Rights Reserved.</copyright>
		
		
			<lastBuildDate>Sun, 09 Aug 2026 00:00:00 +0700</lastBuildDate>
		
			<atom:link href="https://www.phamduytung.com/tags/airllm/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>AirLLM: Nhét một mô hình 2.8 nghìn tỷ tham số qua khe cửa 4GB VRAM</title>
				<link>https://www.phamduytung.com/blog/2026-08-09-airllm-layer-streaming-4gb-vram/</link>
				<pubDate>Sun, 09 Aug 2026 00:00:00 +0700</pubDate>
				<guid>https://www.phamduytung.com/blog/2026-08-09-airllm-layer-streaming-4gb-vram/</guid>
				<description>AirLLM không làm mô hình nhỏ đi. Nó biến GPU thành một căn phòng nhỏ, chỉ đưa từng layer hoặc từng expert vào tính toán rồi lập tức đẩy ra ngoài. Đổi lại, VRAM giảm cực mạnh nhưng disk I/O và latency trở thành cái giá phải trả.</description>
			</item>
	</channel>
</rss>
