<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>CUDAStream on Ringi's Log</title><link>https://lilinji.github.io/tags/cudastream/</link><description>Recent content in CUDAStream on Ringi's Log</description><generator>Hugo -- 0.152.2</generator><language>zh-cn</language><lastBuildDate>Mon, 31 Aug 2026 00:00:00 +0800</lastBuildDate><atom:link href="https://lilinji.github.io/tags/cudastream/index.xml" rel="self" type="application/rss+xml"/><item><title>第07讲：PyTorch 异步执行流——CUDA Stream 与 CUDA Graph 原理</title><link>https://lilinji.github.io/2026/08/%E7%AC%AC07%E8%AE%B2pytorch-%E5%BC%82%E6%AD%A5%E6%89%A7%E8%A1%8C%E6%B5%81cuda-stream-%E4%B8%8E-cuda-graph-%E5%8E%9F%E7%90%86/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0800</pubDate><guid>https://lilinji.github.io/2026/08/%E7%AC%AC07%E8%AE%B2pytorch-%E5%BC%82%E6%AD%A5%E6%89%A7%E8%A1%8C%E6%B5%81cuda-stream-%E4%B8%8E-cuda-graph-%E5%8E%9F%E7%90%86/</guid><description>深入剖析 GPU 异步执行第一性原理：Host-Device 生产者消费者队列、Default Stream 与多流并发调度、Event 流间同步屏障、record_stream 显存踩踏暗礁、CPU Launch Overhead 瓶颈剖析，以及 CUDA Graph 录制回放与 vLLM 多桶图优化工业级系统实战。</description></item></channel></rss>