Using llm-d to Serve Large Models @RedHatOpen
Using llm-d to Serve Large Models  @RedHatOpen
Uploaded March 2026 | Updated September 2026, 9 minutes ago
AMD and Kubernetes communities collaborate on the Gateway API Inference extension to intelligently route requests, making large model serving more efficient through disaggregated prefill and decoding.
Using llm-d to Serve Large ModelsDeferred 2FA for Packagers Proposal - 2026-08-26Mining Issued Common Criteria and FIPS 140-2 Certificates - Red Hat Research Days 2021Starting up Containers Super Fast With Lazy Pulling of ImagesCommunity Central: Exploring DevOps AutomationCommunity Central: Kubernetes + Kubevirt + WindowsAgent has entered the chat - 2026-09-09Triton on AMD GPUsUnikernel  - Red Hat Research Days US 2020BuildKit: Intro to the Architecture of a Modern Build FrameworkKubernetes Service Selection Optimization & Fine-grained Network Telemetry in Programmable SwitchesThe future of Bugzilla Part 2 - 2026-07-29
Red Hat Open |

Using llm-d to Serve Large Models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER