Skip to content
Β 
Β 

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

1,621 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

EzRouter Dashboard

EzRouter - FREE AI Router & Token Saver

Never stop coding. Save 20-40% tokens with RTK + auto-fallback to FREE & cheap AI models.

Connect All AI Code Tools (Claude Code, Cursor, Antigravity, Copilot, Codex, Gemini, OpenCode, Cline, OpenClaw...) to 40+ AI Providers & 100+ Models.

GitHub stars License

Originally forked from decolua/9router β€” upstream attribution kept per MIT license.

πŸš€ Quick Start β€’ πŸ’‘ Features β€’ πŸ“– Setup


πŸ”„ How It Works

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Your CLI   β”‚  (Claude Code, Codex, OpenClaw, Cursor, Cline...)
β”‚   Tool      β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
       β”‚ http://localhost:20128/v1
       ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚           EzRouter (Smart Router)           β”‚
β”‚  β€’ RTK Token Saver (cut tool_result tokens) β”‚
β”‚  β€’ Format translation (OpenAI ↔ Claude)     β”‚
β”‚  β€’ Quota tracking                           β”‚
β”‚  β€’ Auto token refresh                       β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β”œβ”€β†’ [Tier 1: SUBSCRIPTION] Claude Code, Codex, GitHub Copilot
       β”‚   ↓ quota exhausted
       β”œβ”€β†’ [Tier 2: CHEAP] GLM ($0.6/1M), MiniMax ($0.2/1M)
       β”‚   ↓ budget limit
       └─→ [Tier 3: FREE] Kiro, OpenCode Free, Vertex ($300 credits)

Result: Never stop coding, minimal cost + 20-40% token savings via RTK

⚑ Quick Start

1. Install from source:

git clone https://github.com/golamrabbi696/EzRouter.git
cd EzRouter
cp .env.example .env
npm install
npm run build
PORT=20128 HOSTNAME=0.0.0.0 NEXT_PUBLIC_BASE_URL=http://localhost:20128 npm run start

πŸŽ‰ Dashboard opens at http://localhost:20128


πŸ› οΈ Supported CLI Tools

EzRouter works seamlessly with all major AI coding tools:

Claude Code
Claude-Code
OpenClaw
OpenClaw
Codex
Codex
OpenCode
OpenCode
Cursor
Cursor
Antigravity
Antigravity
Cline
Cline
Continue
Continue
Droid
Droid
Roo
Roo
Copilot
Copilot
Kilo Code
Kilo Code
OpenDesign
OpenDesign
jcode
jcode
Grok Build
Grok Build
Devin CLI
Devin CLI
DeepSeek TUI
DeepSeek TUI
Qwen Code
Qwen Code

🌐 Supported Providers

πŸ” OAuth Providers

Claude Code
Claude-Code
Antigravity
Antigravity
Codex
Codex
GitHub
GitHub
Cursor
Cursor
Kimchi
Kimchi

πŸ†“ Free Providers

Kiro
Kiro AI
Claude 4.5 + GLM-5 + MiniMax
50 credits/month free
OpenCode Free
OpenCode Free
No auth β€’ Auto-fetch models
Free (model list varies)
Vertex AI
Vertex AI
Gemini 3 Pro + GLM-5 + DeepSeek
$300 credits free

Note: iFlow, Qwen Code and Gemini CLI free tiers were discontinued in 2026. Use Kiro / OpenCode Free / Vertex instead.

Kiro AI moved to a paid model in Sep 2025 β€” the free tier is now capped at 50 credits/month (plus 500 trial credits for new accounts in the first 30 days). Paid tiers: Pro $20/mo (1,000 credits), Pro+ $40/mo (2,000), Pro Max $100/mo (5,000), Power $200/mo (10,000). OpenCode Free model list fluctuates over time (some models free only for limited promos) β€” subject to change without notice. Vertex AI: the $300 free credit for new GCP accounts is still valid, but since Mar 2026 the Gemini API endpoint no longer consumes these credits β€” call the Vertex AI Studio endpoint instead.

πŸ”‘ API Key Providers (40+)

OpenRouter
OpenRouter
GLM
GLM
Kimi
Kimi
MiniMax
MiniMax
OpenAI
OpenAI
Anthropic
Anthropic
Gemini
Gemini
DeepSeek
DeepSeek
Groq
Groq
xAI
xAI
Mistral
Mistral
Perplexity
Perplexity
Together
Together AI
Fireworks
Fireworks
Cerebras
Cerebras
Cohere
Cohere
NVIDIA
NVIDIA
SiliconFlow
SiliconFlow

...and 20+ more providers including Nebius, Chutes, Hyperbolic, and custom OpenAI/Anthropic compatible endpoints

🏠 Self-hosted Providers

For speech and embeddings served from your own machine β€” whisper.cpp, faster-whisper, Speaches, Kokoro-FastAPI, openedai-speech, llama.cpp/llama-server, vLLM, Infinity, text-embeddings-inference, or anything else that speaks the OpenAI shape.

Provider Endpoint used Typical server
Self-hosted STT /v1/audio/transcriptions whisper.cpp, faster-whisper
Self-hosted TTS /v1/audio/speech Kokoro-FastAPI, openedai-speech
Self-hosted Embedding /v1/embeddings llama-server, vLLM, Infinity

Every other speech provider is a named cloud service with a fixed endpoint. These three read their address from each connection, so one provider can front several machines and load-balance across them like any other.

Set it on the connection as providerSpecificData.baseUrl:

Provider Give it Result
Self-hosted STT the full URL β€” http://host:8080/v1/audio/transcriptions used as-is
Self-hosted TTS the server root β€” http://host:8880 + /v1/audio/speech
Self-hosted Embedding the OpenAI base, /v1 included β€” http://host:8080/v1 + /embeddings

Mind the /v1 on embeddings. The adapter appends /embeddings, so http://host:8080 resolves to http://host:8080/embeddings and misses the OpenAI route β€” llama-server answers 501. Give it the same base URL an OpenAI client would use. A full .../v1/embeddings is also accepted, so a value pasted from a curl example works too.

The API key is not checked by most local servers, but the field must be non-empty: it is what gives the connection a credentials record, and baseUrl lives there. Any placeholder works.

Self-hosted Embedding has no cloud fallback by design β€” a connection saved without a baseUrl is reported as a configuration error rather than quietly falling back to api.openai.com, which would send your input text and API key to a third party under a provider named "Self-hosted".


πŸ’‘ Key Features

Feature What It Does Why It Matters
πŸš€ RTK Token Saver (RTK ⭐40K) Compress tool outputs (git diff, grep, ls, tree...) before sending to LLM Save 20-40% input tokens per request
🧠 Headroom Token Saver (Headroom) Optional external /v1/compress proxy before provider routing Save more context tokens without changing clients
πŸͺ¨ Caveman Mode (Caveman ⭐52K) Inject caveman-speak prompt β†’ LLM replies terse, technical substance preserved Save up to 65% output tokens
🐴 Ponytail (Ponytail) Inject "lazy senior dev" prompt β†’ LLM writes minimal, YAGNI-first code (Lite/Full/Ultra) Fewer output tokens, less refactoring
🎯 Smart 3-Tier Fallback Auto-route: Subscription β†’ Cheap β†’ Free Never stop coding, zero downtime
πŸ“Š Real-Time Quota Tracking Live token count + reset countdown Maximize subscription value
πŸ”„ Format Translation OpenAI ↔ Claude ↔ Gemini ↔ Cursor ↔ Kiro ↔ Vertex Works with any CLI tool
πŸ‘₯ Multi-Account Support Multiple accounts per provider Load balancing + redundancy
πŸ”„ Auto Token Refresh OAuth tokens refresh automatically No manual re-login needed
🎨 Custom Combos Create unlimited model combinations Tailor fallback to your needs
πŸ“ Request Logging Debug mode with full request/response logs Troubleshoot issues easily
πŸ’Ύ Cloud Sync Sync config across devices Same setup everywhere
πŸ“Š Usage Analytics Track tokens, cost, trends over time Optimize spending
🌐 Deploy Anywhere Localhost, VPS, Docker, Cloudflare Workers Flexible deployment options

Set X-9Router-Token-Saver: off to bypass all token savers for one chat request. (Header name unchanged from upstream.)


πŸ› οΈ Tech Stack

  • Runtime: Node.js 20+
  • Framework: Next.js 16
  • UI: React 19 + Tailwind CSS 4
  • Database: SQLite (better-sqlite3 / node:sqlite / sql.js fallback)
  • Streaming: Server-Sent Events (SSE)
  • Auth: OAuth 2.0 (PKCE) + JWT + API Keys

πŸ“ API Reference

Chat Completions

POST http://localhost:20128/v1/chat/completions
Authorization: Bearer your-api-key
Content-Type: application/json

{
  "model": "cc/claude-opus-4-6",
  "messages": [
    {"role": "user", "content": "Write a function to..."}
  ],
  "stream": true
}

List Models

GET http://localhost:20128/v1/models
Authorization: Bearer your-api-key

β†’ Returns all models + combos in OpenAI format

πŸ“§ Support


πŸ‘₯ Contributors

Thanks to all contributors who helped make EzRouter better!

Contributors


πŸ™ Acknowledgments

Built on the shoulders of giants:

  • CLIProxyAPI β€” original Go implementation that inspired this JavaScript port.
  • RTK Stars β€” Rust token-saver. EzRouter ports its compression pipeline to JS β†’ βˆ’20-40% input tokens on every request.
  • Caveman Stars by @JuliusBrussee β€” viral "why use many token when few token do trick". EzRouter adapts its prompt β†’ βˆ’65% output tokens.
  • Ponytail Stars by @DietrichGebert β€” "lazy senior dev" skill. EzRouter injects its YAGNI-first ladder β†’ fewer tokens, less code, shorter diffs.
  • decolua/9router β€” the upstream project this fork is based on. All core functionality originates there.

Huge thanks to these authors β€” without their work, EzRouter's token-saving features wouldn't exist. ⭐ them on GitHub!


πŸ“„ License

MIT License - see LICENSE for details.


Built with ❀️ for developers who code 24/7
Maintained by Golam Rabbi

About

Unlimited FREE AI coding. Connect Claude Code, Codex, Cursor, Cline, Copilot, Antigravity to FREE Claude/GPT/Gemini via 40+ providers. Auto-fallback, RTK -40% tokens, never hit limits.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages