Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
Uh oh!
There was an error while loading.
Please reload this page
.
AMD-Ecosystem
/
llama.cpp
Public
forked from
ggml-org/llama.cpp
Notifications
You must be signed in to change notification settings
Fork
8
Star
18
Code
Issues
1
Pull requests
16
Discussions
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Discussions
Actions
Projects
Security and quality
Insights
Actions: AMD-Ecosystem/llama.cpp
Actions
All workflows
Workflows
Code Style Checker
Code Style Checker
.github/workflows/build-cann.yml
.github/workflows/build-cann.yml
AI review (issues)
AI review (issues)
Build Actions Cache
Build Actions Cache
Build gfx11 + ROCm
Build gfx11 + ROCm
CI
CI
CI (3rd-party)
CI (3rd-party)
CI (android)
CI (android)
CI (cross)
CI (cross)
CI (CUDA, windows)
CI (CUDA, windows)
CI (msys)
CI (msys)
Show more workflows...
Management
Caches
Code Style Checker
Code Style Checker
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Show workflow options
Create status badge
Create status badge
Loading
Uh oh!
There was an error while loading.
Please reload this page
.
code-style.yml
will be ignored since log searching is not yet available
190 workflow runs
190 workflow runs
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#123:
Pull request
#59
synchronize by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#121:
Pull request
#59
synchronize by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#120:
Pull request
#59
synchronize by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
mmq: compacted MoE tiling for RDNA3.5 (default on)
Code Style Checker
#115:
Pull request
#63
synchronize by
roberteg16
14s
rogarcia.mmq-moe-compact-tiling
rogarcia.mmq-moe-compact-tiling
14s
View #63
View workflow file
mmq: compacted MoE tiling for RDNA3.5 (default on)
Code Style Checker
#114:
Pull request
#63
synchronize by
roberteg16
55s
rogarcia.mmq-moe-compact-tiling
rogarcia.mmq-moe-compact-tiling
55s
View #63
View workflow file
mmq: compacted MoE tiling for RDNA3.5 (default on)
Code Style Checker
#113:
Pull request
#63
synchronize by
roberteg16
43s
rogarcia.mmq-moe-compact-tiling
rogarcia.mmq-moe-compact-tiling
43s
View #63
View workflow file
mmq: compacted MoE tiling for RDNA3.5 (default on)
Code Style Checker
#112:
Pull request
#63
opened by
roberteg16
26s
rogarcia.mmq-moe-compact-tiling
rogarcia.mmq-moe-compact-tiling
26s
View #63
View workflow file
RDNA3.5 (gfx11 / gfx1151) MMQ prefill optimizations
Code Style Checker
#111:
Pull request
#32
synchronize by
liangliangchang
35s
liangliangchang:lichang.mmq_opt
liangliangchang:lichang.mmq_opt
35s
View #32
View workflow file
RDNA3.5 (gfx11 / gfx1151) MMQ prefill optimizations
Code Style Checker
#110:
Pull request
#32
synchronize by
liangliangchang
25s
liangliangchang:lichang.mmq_opt
liangliangchang:lichang.mmq_opt
25s
View #32
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#109:
Pull request
#59
synchronize by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#108:
Pull request
#59
synchronize by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
RDNA3.5 (gfx11 / gfx1151) MMQ prefill optimizations
Code Style Checker
#107:
Pull request
#32
synchronize by
liangliangchang
17s
liangliangchang:lichang.mmq_opt
liangliangchang:lichang.mmq_opt
17s
View #32
View workflow file
Merge pull request #56 from ROCm/rogarcia.enable-by-default-ggml_cuda…
Code Style Checker
#106:
Commit
6e9f948
pushed by
mgehre-amd
56s
gfx11
gfx11
56s
View workflow file
tests: add MoE MMQ benchmark with routing-distribution generator
Code Style Checker
#105:
Pull request
#62
opened by
roberteg16
1d 0h 0m 1s
rogarcia.moe-mmq-benchmark
rogarcia.moe-mmq-benchmark
1d 0h 0m 1s
View #62
View workflow file
ggml-cuda: add dequant-float matvec (mmvdq) for Q4_K/Q5_K/Q6_K on RDN…
Code Style Checker
#104:
Pull request
#61
opened by
Annieren
20s
annier.mmvq-to-mmvdq
annier.mmvq-to-mmvdq
20s
View #61
View workflow file
Gfx11
Code Style Checker
#103:
Pull request
#60
opened by
Annieren
14s
gfx11
gfx11
14s
View #60
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#102:
Pull request
#59
synchronize by
jeffli-xilinx
24s
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
24s
View #59
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#101:
Pull request
#59
synchronize by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#100:
Pull request
#59
synchronize by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
Code Style Checker
#99:
Pull request
#59
opened by
jeffli-xilinx
Action required
jeffli-xilinx:rdna3_5-kv-fusion
jeffli-xilinx:rdna3_5-kv-fusion
Action required
View #59
View workflow file
ggml-cuda: GEMM weight row padding + one-time K-padded f16 dequant for prefill
Code Style Checker
#97:
Pull request
#57
synchronize by
roberteg16
1m 4s
rogarcia.cuda-gemm-pad-prefill-dequant
rogarcia.cuda-gemm-pad-prefill-dequant
1m 4s
View #57
View workflow file
ggml-cuda: GEMM weight row padding + one-time K-padded f16 dequant for prefill
Code Style Checker
#96:
Pull request
#57
opened by
roberteg16
32s
rogarcia.cuda-gemm-pad-prefill-dequant
rogarcia.cuda-gemm-pad-prefill-dequant
32s
View #57
View workflow file
ggml-cuda: enable GGML_CUDA_GRAPH_OPT by default on RDNA3.5
Code Style Checker
#95:
Pull request
#56
opened by
roberteg16
46s
rogarcia.enable-by-default-ggml_cuda_graph_opt-on
rogarcia.enable-by-default-ggml_cuda_graph_opt-on
46s
View #56
View workflow file
Merge pull request #53 from ROCm/rogarcia.roofline-moe-tokens-per-expert
Code Style Checker
#94:
Commit
0941c4c
pushed by
mgehre-amd
42s
gfx11
gfx11
42s
View #60
View workflow file
CUDA: fused chunked gated_delta_net kernel (RDNA3.5)
Code Style Checker
#91:
Pull request
#54
synchronize by
roberteg16
22s
rogarcia.gdn-shared-kq
rogarcia.gdn-shared-kq
22s
View #54
View workflow file
Previous
1
2
3
4
5
6
7
8
Next
You can’t perform that action at this time.