Research Wiki
23 pages across 3 groups.
Docs (4)
Freetoken (1)
Top level (18)
- FreeToken CUDA → ROCm Adaptation — Overview & Surface Map
- 01 — CUDA→HIP Porting Toolchain (HIPify, torch ROCm extension, hipcc)
- 02 — GGUF Dequant Kernels: the llama.cpp ggml-hip Reference
- 03 — Triton Kernels on AMD (HIP backend + FP8/format limits)
- 04 — Attention Backends on AMD (FlashAttention / FlashInfer / AOTriton)
- 05 — RCCL: the NCCL Drop-in for pynccl
- 06 — NVFP4 / MXFP4 / MXFP8 / FP8 → bf16/fp16 fallback
- 07 — Target Hardware gfx906 (Radeon VII / Instinct MI50): Capabilities + ROCm Status
- 08 — SGLang / FlashInfer / AITER ROCm Ecosystem Status
- 09 — Donato Capitella (kyuz0): gfx906 Toolboxes and the Sources He Used
- 10 — The Full Quantization Matrix: IQ1_S Through Q8_0 (Removing the Excerpt's Limits)
- 11 — Arch Linux / AUR: The Packaged ROCm Stack (Build Discipline Without Local Heroics)
- 12 — Nobara / Fedora Current Track: Triton Recompilation + the mxxm gfx906 Backend
- 13 — Scarcity as a Compiler: How Qwen, DeepSeek, and GLM Engineered Around Compute Limits
- 14 — Whitepaper Catalog: The Papers the Scarcity Workarounds Produced
- DU1 Code Analysis — FreeToken CUDA Surface Audit (Stage 1)
- DU1 Decision Tree — FreeToken CUDA→ROCm/gfx906 Port (Stage 2)
- DU1 File-Impact Map — Files Affected by the Port (Stage 1 output, tree-feeding)