Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
Uh oh!
There was an error while loading.
Please reload this page
.
PrismML-Eng
/
llama.cpp
Public
forked from
ggml-org/llama.cpp
Notifications
You must be signed in to change notification settings
Fork
88
Star
463
Code
Issues
17
Pull requests
18
Discussions
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Discussions
Actions
Projects
Security and quality
Insights
Actions: PrismML-Eng/llama.cpp
Actions
All workflows
Workflows
Release (Prism)
Release (Prism)
Server
Server
Build relocatable cmake package
Build relocatable cmake package
Check Pre-Tokenizer Hashes
Check Pre-Tokenizer Hashes
Check vendor
Check vendor
Code Style Checker
Code Style Checker
CodeQL
CodeQL
Copilot
Copilot
Copilot cloud agent
Copilot cloud agent
Copilot code review
Copilot code review
Copilot Setup Steps
Copilot Setup Steps
Show more workflows...
Management
Caches
Server
Server
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Show workflow options
Create status badge
Create status badge
Loading
Uh oh!
There was an error while loading.
Please reload this page
.
server.yml
will be ignored since log searching is not yet available
125 workflow runs
125 workflow runs
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
llama/ggml: Hadamard weight-fold runtime (prism.hadamard GGUF contract, FWHT to 2048)
Server
#179:
Pull request
#119
opened by
bri-prism
20m 55s
hadamard-gguf-contract
hadamard-gguf-contract
20m 55s
View #119
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#178:
Pull request
#116
synchronize by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#177:
Pull request
#116
synchronize by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#176:
Pull request
#116
synchronize by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#175:
Pull request
#116
synchronize by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
metal: fix the recurrent-state snapshot ring's CPU fallback and wide-row set_rows
Server
#174:
Pull request
#118
synchronize by
bri-prism
19m 20s
fix/metal-im2col-dispatch-wide-set-rows
fix/metal-im2col-dispatch-wide-set-rows
19m 20s
View #118
View workflow file
metal: fix the recurrent-state snapshot ring's CPU fallback and wide-row set_rows
Server
#173:
Pull request
#118
synchronize by
bri-prism
3m 34s
fix/metal-im2col-dispatch-wide-set-rows
fix/metal-im2col-dispatch-wide-set-rows
3m 34s
View #118
View workflow file
metal: fix the recurrent-state snapshot ring's CPU fallback and wide-row set_rows
Server
#172:
Pull request
#118
opened by
bri-prism
19m 8s
fix/metal-im2col-dispatch-wide-set-rows
fix/metal-im2col-dispatch-wide-set-rows
19m 8s
View #118
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#171:
Pull request
#116
synchronize by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#169:
Pull request
#116
synchronize by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#168:
Pull request
#116
synchronize by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
cuda: AMD RDNA4 Q1_0/Q2_0 — HIP-path vec_dots (+37%/+16% decode), opt-in quant dedup, opt-in hipBLASLt prefill routes
Server
#167:
Pull request
#116
opened by
The-Monk
Action required
The-Monk:rdna4-q1q2
The-Monk:rdna4-q1q2
Action required
View #116
View workflow file
prism-v6: dspark Path A, converge on upstream dflash with log-SNR conditioning and drafter-own embeddings
Server
#166:
Pull request
#114
synchronize by
bri-prism
21m 31s
prism-v6-dspark-dflash
prism-v6-dspark-dflash
21m 31s
View #114
View workflow file
prism-v6: complete Q2_0_g128 switch coverage in ops.cpp (fixes -Werror=switch CI)
Server
#165:
Pull request
#115
opened by
bri-prism
18m 11s
prism-v6-g128-switch-coverage
prism-v6-g128-switch-coverage
18m 11s
View #115
View workflow file
prism-v6: fix Q2_0_g128 ARM link failure and Metal shader compile
Server
#164:
Pull request
#113
synchronize by
khosravipasha
21m 57s
prism-v6-g128-arm-metal-fixes
prism-v6-g128-arm-metal-fixes
21m 57s
View #113
View workflow file
prism-v6: dspark Path A, converge on upstream dflash with log-SNR conditioning and drafter-own embeddings
Server
#163:
Pull request
#114
synchronize by
khosravipasha
21m 39s
prism-v6-dspark-dflash
prism-v6-dspark-dflash
21m 39s
View #114
View workflow file
prism-v6: fix Q2_0_g128 ARM link failure and Metal shader compile
Server
#162:
Pull request
#113
synchronize by
khosravipasha
7m 28s
prism-v6-g128-arm-metal-fixes
prism-v6-g128-arm-metal-fixes
7m 28s
View #113
View workflow file
cuda: MoE prefill — fused SwiGLU epilogue and fused weighted reduction
Server
#161:
Pull request
#107
synchronize by
bri-prism
24m 10s
perf/moe-prefill-cda
perf/moe-prefill-cda
24m 10s
View #107
View workflow file
cuda: enable the Hopper wgmma Q1_0/Q2_0 prefill path by default
Server
#160:
Pull request
#106
synchronize by
bri-prism
23m 39s
perf/hopper-q1-default-on
perf/hopper-q1-default-on
23m 39s
View #106
View workflow file
cuda: MoE prefill — fused SwiGLU epilogue and fused weighted reduction
Server
#159:
Pull request
#107
synchronize by
bri-prism
19m 18s
perf/moe-prefill-cda
perf/moe-prefill-cda
19m 18s
View #107
View workflow file
cuda: MoE prefill — fused SwiGLU epilogue and fused weighted reduction
Server
#158:
Pull request
#107
opened by
bri-prism
16m 49s
perf/moe-prefill-cda
perf/moe-prefill-cda
16m 49s
View #107
View workflow file
cuda: enable the Hopper wgmma Q1_0/Q2_0 prefill path by default
Server
#157:
Pull request
#106
opened by
bri-prism
20m 8s
perf/hopper-q1-default-on
perf/hopper-q1-default-on
20m 8s
View #106
View workflow file
speculative: window dspark drafter staging to its trained position range
Server
#156:
Pull request
#98
synchronize by
thadreber-web
Action required
thadreber-web:fix/dspark-windowed-staging
thadreber-web:fix/dspark-windowed-staging
Action required
View #98
View workflow file
server: clear slot state on task cancellation
Server
#155:
Pull request
#103
synchronize by
Copilot
AI
Action required
copilot/fix-cuda-illegal-memory-access
copilot/fix-cuda-illegal-memory-access
Action required
View #103
View workflow file
mtmd: add lanczos resize method [no release] (#26341)
Server
#154:
Commit
5f55650
pushed by
khosravipasha
17m 26s
master
master
17m 26s
View workflow file
Previous
1
2
3
4
5
Next
You can’t perform that action at this time.