Inference · Python
ProppaTP
CPU-stage prototypeProppaTP explores prompt preprocessing and tensor-parallel LLM decoding on consumer GPUs. The published project is a CPU-stage prototype; GPU execution, nvFP4 loading, and the prefill/decode split are development targets, not demonstrated throughput results.
github.com/AStarStarship/… →Specification
| Language | Python (PyTorch) |
|---|---|
| Target rig | 1x RTX 5070 Ti (prefill) + 2x RTX 5060 Ti (TP=2 decode) |
| Target interconnect | PCIe 5.0 x16 per card (no NVLink) |
| Quantization target | nvFP4; model loading is pending |
Design & implementation
Planned prefill / decode split
The target architecture assigns prompt encoding to one GPU and tensor-parallel decode to a pair of GPUs. The public checklist still marks this GPU split as unfinished.
Row-parallel over PCIe
The design uses row-parallel linear operations to limit communication over PCIe without NVLink. TP=2 shard execution and GPU all-reduce still need hardware validation.
Cache development
Paged KV caching is a development target. GPU cache behavior and memory use need measurement before the project can claim a throughput improvement.
Target model allocation
The README outlines separate auxiliary and decode models. Their nvFP4 loading, GPU placement, and memory budgets are targets to validate, not a working multi-model deployment demonstrated here.
Related