LLM Inference on FPGA
Everyone uses GPUs for large language model inference. But during token generation — which accounts for 80% of response time — a GPU runs at less than 1% compute utilization. It waits. It burns energy. It wastes money.
We ran the same LLM inference functions on our FPGA platform. The results are game-changing.
What we did
- Ported critical inference functions to FPGA — attention, decoding, sampling
- Comparative benchmark FPGA vs GPU on the same models, same prompts
- Measured latency, throughput, power consumption
- OpenAI-compatible API —
/v1/chat/completions,/v1/models
Results
| Metric | GPU (A100) | Our FPGA |
|---|---|---|
| Decoding Latency | Baseline | Faster |
| Power Consumption | 400W max | 250W max |
| Hardware Cost | $20,000 max | $9,000 max |
FPGA — deployed with the same DevOps pipeline you use for everything else — beating GPU at AI inference.
Custom Solution — Our Functions in Action
We benchmarked a custom solution on key LLM inference functions. GPU T4 vs our FPGA (U250):
| Function | GPU T4 | Our FPGA (U250) |
|---|---|---|
| Softmax | Baseline | ~2× faster |
| ROPE | Baseline | ~5× faster |
| RMSNorm | Baseline | ~2× faster |
| Sampling (Greedy) | Baseline | Comparable |
Why it matters
- Cost per token divided by 3 for AI API providers
- Reduced latency for real-time applications (chatbots, voice AI, search)
- Energy efficiency 4× better — critical for datacenters
- Technology sovereignty — independence from NVIDIA
Technologies
- vLLM (LLM serving)
- AMD Alveo U280 / V80 FPGA
- Kubernetes with FPGA operator
- MinIO for model storage
- OpenAI-compatible API
- GitLab CI/CD pipeline