MI300X MoE Training
075 - TRL + Custom megablocks-hip fork gets to step 0 at least Sample Blog - GPT2 training: https://rocm.blogs.amd.com/artificial-intelligence/megablocks/README.
Last 10 notes on š llm-tracker
075 - TRL + Custom megablocks-hip fork gets to step 0 at least Sample Blog - GPT2 training: https://rocm.blogs.amd.com/artificial-intelligence/megablocks/README.
llama.cpp llama2-7b 600W ⯠build/bin/llama-bench -m /models/llm/gguf/llama-2-7b.Q4_0.gguf -fa 1 ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no ggml_cuda_init: found 1 CUDA devices: Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.
https://www.notion.so/Ascend_Doc-2180dc3bf51680989af4cd4eee46acdd Training (ms-swift) https://swift.readthedocs.io/en/latest/BestPractices/NPU-support.
Those looking for my testing code: https://github.com/lhl/strix-halo-testing I will probably consolidate and redirect to the Github repo at some point.
For best price/perf, Dual Socket EPYC ROME is probably the way to go. If you have the cash: Cheap dual core 9004 chips. The going rate for a 9334 QS/ES chip is 600atm,soyoucouldgetapairfor1200 and should give you about 400GB/s For dual socket youāre probably looking at a Gigabyte MZ73-LM1 or AsRock Rack TURIN2D16-2T - looks like those arenāt going for about $1000-1300 Itās about $1500-1800 forā¦
Practical mlabonneās LLM Course https://github.com/mlabonne/llm-course Mastering LLMs https://hamel.dev/blog/posts/course/ Evals https://www.youtube.com/playlist?list=PLgIaq8VgndJvt-HKMHPXehyJNNXQsAVHD https://hamel.
FP16 --- Scores for Model: shisa-ai/shisa-v2-llama3.1-405b --- Category gpt-4-0613 gpt-4-turbo gpt-4.1-2025-04-14 gpt-4.1-mini-2025-04-14 gpt-4o ----------- ------------ ------------- -------------------- ------------------------- -------- coding 9.
As of August 2023, AMDās ROCm GPU compute software stack is available for Linux or Windows. Itās best to check the latest docs for information: https://rocm.
NOTE: This document tries to avoid using the term āperformanceā since in ML research the term performance typically refers to measuring model quality/capabilities.