Hacker News AI · 2026/10/7 13:15:44
开源 rGPU:实现 PyTorch 张量远程 GPU 执行与 CUDA 兼容层透传
rGPU 是一款开源工具,允许应用程序在本地运行而将计算任务卸载至远程 NVIDIA GPU。它提供两种集成路径:一是作为 PyTorch 设备直接支持 `rgpu` 后端,二是通过 CUDA Shim 拦截底层 API 以兼容现有二进制程序。该方案旨在解决本地算力不足或资源隔离场景下的远程训练与推理需求。
报道全文原始报道全文
本文目录4 个章节
rGPU
rGPU runs GPU work on a remote NVIDIA machine while the application stays on the client. It currently offers two paths:
The PyTorch device is the simpler integration. The CUDA shim covers existing binaries but has a larger compatibility surface.
Documentation
The Fumadocs site in website/ is the product documentation:
- Quickstart
- Training
- nanoGPT example
- CUDA shim
- Operations
- Configuration reference
- Performance
- Troubleshooting
Engineering records and experiments are indexed in docs/README.md.
Quick start: PyTorch device
Install rGPU with pip install rgpu, or pip install -e ./python from this
checkout, then follow the quickstart to
deploy the server. Save this as smoke.py in your workload directory:
import torch
import rgpu
x = torch.ones(4, device="rgpu")
print((x * 2).sum().item()) # 8.0Run it in the environment where rGPU is installed, using your server's SSH destination and options:
rgpu-run --host user@gpu-host --ssh-port 2222 -i ~/.ssh/gpu_key \
python smoke.pyThe program selects the device; rgpu-run opens the tunnel and configures the
connection. The expected output is 8.0.
For existing Linux CUDA programs, follow the
CUDA shim guide, starting with
./scripts/build_client.sh.
Neither protocol authenticates or encrypts connections. Keep rgpu-opserver
on its default localhost bind and use SSH. The CUDA server listens on all IPv4
interfaces: restrict port 9713 with host/cloud firewall rules before starting
it, even when using an SSH tunnel. See deployment.
Development
# C++ client and fake-driver tests
./scripts/build_client.sh
# Python tests
python -m pip install -e './python[test]'
python -m pytest python/tests
# Static documentation
npm --prefix website ci
npm --prefix website run buildSee scripts/README.md for the remaining build, cloud,
and hardware commands. Generated C++ is committed; its policy and regeneration
steps are in codegen/README.md.
Repository map
| Path | Purpose |
|---|---|
client/ | CUDA client shims and transport |
server/ | CUDA server and dispatch |
common/ | Shared protocol and generated API metadata |
python/ | PyTorch device and launcher |
tests/ | C++, Python, CUDA, and hardware checks |
codegen/ | CUDA header parser and source generators |
website/ | Fumadocs product documentation |
docs/ | Design records, measurements, and experiment reports |
jax/ | Experimental JAX work; not a supported product path |
scripts/ | Build, deployment, cloud, and test helpers |
skills/ | Installable agent guidance for using rGPU |
Historical implementation notes and experimental results are indexed in
docs/README.md.
作者发帖说明(HN):
rGPU runs GPU work on a remote NVIDIA machine while the application stays on the client.