{"version":1,"type":"story","url":"https://digestai.news/story/github-releases-kernelopt-for-agentic-gpu-kernel-optimization","json":"https://digestai.news/story/github-releases-kernelopt-for-agentic-gpu-kernel-optimization.json","markdown":"https://digestai.news/story/github-releases-kernelopt-for-agentic-gpu-kernel-optimization.md","slug":"github-releases-kernelopt-for-agentic-gpu-kernel-optimization","headline":"GitHub releases KernelOPT for agentic GPU kernel optimization","summary":"GitHub user **giveen** released **KernelOPT**, an open-source tool that automatically optimizes GPU kernels using agentic search. It works on CUDA, HIP/ROCm, and frameworks like PyTorch (Triton), ninfer, and llama.cpp. KernelOPT profiles kernels, generates optimized versions, and verifies correctness before applying changes in an isolated git worktree—preserving the original codebase. The tool uses LLM agents (Planner, Executor, Summarizer) guided by profiling data (nsys/ncu or Graphsignal) and a multi-stage validation loop to ensure speedups are real and physically plausible.\n\nKey features include **five gates** for correctness (compile, multi-seed, model-level, performance, and shape correctness) and support for **cloud-based LLM optimization** (e.g., OpenAI, OpenRouter) to avoid GPU contention. Users can run it on NVIDIA/AMD GPUs with CUDA/ROCm toolkits, Rust, Python, and CMake. The project is licensed under **Apache 2.0** and builds on research from arXiv:2609.30059. Documentation and setup instructions are available in the repo.","keyPoints":["KernelOPT automates GPU kernel optimization via agentic search, supporting CUDA, HIP/ROCm, and frameworks like PyTorch, ninfer, and llama.cpp","Uses LLM agents (Planner/Executor/Summarizer) with profiling tools (nsys/ncu, Graphsignal) to generate, verify, and apply optimized kernels without touching the original repo","Five correctness gates ensure speedups are measurable, plausible, and reproducible; runs on cloud LLMs to avoid GPU resource conflicts"],"whyItMatters":"KernelOPT could slash GPU compute costs for AI training by optimizing low-level kernels—critical for labs running large models on limited hardware. Its agentic approach may inspire broader automation in hardware optimization.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["GitHub"],"models":[],"people":["Aheli Poddar","Sanskar Prasad","Arindam Samanta","Subha Chakraborty","Vishal Goyal","Rohit Singh Rathaur"]},"firstPublishedAt":"2026-09-29T00:00:00Z","updatedAt":"2026-09-29T00:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"github.com","title":"GitHub - giveen/KernelOPT: Dispatch-aware agentic GPU kernel optimization","url":"https://github.com/giveen/KernelOPT","publishedAt":"2026-09-29T00:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wvwgaz/github_giveenkernelopt_dispatchaware_agentic_gpu/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"GitHub releases KernelOPT for agentic GPU kernel optimization\", 29 September 2026, https://digestai.news/story/github-releases-kernelopt-for-agentic-gpu-kernel-optimization","publisher":"Digest AI","title":"GitHub releases KernelOPT for agentic GPU kernel optimization","datePublished":"2026-09-29T00:00:00Z","url":"https://digestai.news/story/github-releases-kernelopt-for-agentic-gpu-kernel-optimization"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}