LLM Inference on FPGA

Everyone uses GPUs for large language model inference. But during token generation — which accounts for 80% of response time — a GPU runs at less than 1% compute utilization. It waits. It burns energy. It wastes money.

We ran the same LLM inference functions on our FPGA platform. The results are game-changing.

What we did

  • Ported critical inference functions to FPGA — attention, decoding, sampling
  • Comparative benchmark FPGA vs GPU on the same models, same prompts
  • Measured latency, throughput, power consumption
  • OpenAI-compatible API — /v1/chat/completions, /v1/models

Results

MetricGPU (A100)Our FPGA
Decoding LatencyBaselineFaster
Power Consumption400W max250W max
Hardware Cost$20,000 max$9,000 max

FPGA — deployed with the same DevOps pipeline you use for everything else — beating GPU at AI inference.

Custom Solution — Our Functions in Action

We benchmarked a custom solution on key LLM inference functions. GPU T4 vs our FPGA (U250):

FunctionGPU T4Our FPGA (U250)
SoftmaxBaseline~2× faster
ROPEBaseline~5× faster
RMSNormBaseline~2× faster
Sampling (Greedy)BaselineComparable

Why it matters

  • Cost per token divided by 3 for AI API providers
  • Reduced latency for real-time applications (chatbots, voice AI, search)
  • Energy efficiency 4× better — critical for datacenters
  • Technology sovereignty — independence from NVIDIA

Technologies

  • vLLM (LLM serving)
  • AMD Alveo U280 / V80 FPGA
  • Kubernetes with FPGA operator
  • MinIO for model storage
  • OpenAI-compatible API
  • GitLab CI/CD pipeline

Back to projects