Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 0 additions & 1 deletion providers/together-ai/ByteDance-Seed/Seedream-3.0.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@ costs:
modalities:
input:
- text
- image
output:
- image
mode: image
Expand Down
3 changes: 1 addition & 2 deletions providers/together-ai/ByteDance/Seedream-5.0-lite.yaml
Original file line number Diff line number Diff line change
@@ -1,6 +1,5 @@
costs:
- input_cost_per_token: 0.0
output_cost_per_token: 0.0
- output_cost_per_image: 0.035
region: "*"
modalities:
input:
Expand Down
4 changes: 3 additions & 1 deletion providers/together-ai/Qwen/Qwen3-32B.yaml
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
deprecationDate: "2026-07-29"
features:
- function_calling
- system_messages
isDeprecated: true
limits:
context_window: 128000
max_output_tokens: 32768
Expand All @@ -18,7 +20,7 @@ removeParams:
sources:
- https://www.together.ai/models/qwen3-32b
- https://huggingface.co/Qwen/Qwen3-32B
status: active
status: deprecated
supportedModes:
- chat
thinking: true
5 changes: 5 additions & 0 deletions providers/together-ai/Qwen/Qwen3.5-35B-A3B-Lora.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,11 @@ costs:
region: "*"
limits:
context_window: 262144
modalities:
input:
- text
output:
- text
mode: chat
model: Qwen/Qwen3.5-35B-A3B-Lora
provisioning: provisioned
Expand Down
1 change: 1 addition & 0 deletions providers/together-ai/Qwen/Qwen3.6-35B-A3B-Lora.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ provisioning: provisioned
sources:
- https://docs.together.ai/docs/fine-tuning-models
- https://www.together.ai/pricing
status: active
supportedModes:
- chat
thinking: true
3 changes: 1 addition & 2 deletions providers/together-ai/alibaba/happyhorse-1.0-i2v.yaml
Original file line number Diff line number Diff line change
@@ -1,6 +1,5 @@
costs:
- input_cost_per_token: 0.0 # not found in official docs (i2v variant not listed on Together AI pricing/serverless pages)
output_cost_per_token: 0.0 # not found in official docs

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

HappyHorse costs use wrong unit

High Severity

New HappyHorse prices are stored as input_cost_per_request (0.24 for 1.0, 0.14 for 1.1), treating per-second video rates as a flat per-generation charge. Together lists HappyHorse 1.1 T2V at $0.14 per second, and this repo already records 1.0 T2V as output_cost_per_second_1080p: 0.24. Cost estimates would undercount a multi-second clip by roughly the clip length. The same unit error is in happyhorse-1.1-r2v.yaml.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit e362e08. Configure here.

- input_cost_per_request: 0.24 # $0.24 per video generation
region: "*"
modalities:
input:
Expand Down
5 changes: 3 additions & 2 deletions providers/together-ai/alibaba/happyhorse-1.0-r2v.yaml
Original file line number Diff line number Diff line change
@@ -1,6 +1,5 @@
costs:
- input_cost_per_token: 0.0 # not found in official docs (Together AI changelog lists alibaba/happyhorse-1.0-r2v as serverless but no published price)
output_cost_per_token: 0.0 # not found in official docs
- input_cost_per_request: 0.24
region: "*"
modalities:
input:
Expand All @@ -11,6 +10,8 @@ modalities:
mode: video
model: alibaba/happyhorse-1.0-r2v
provisioning: serverless
sources:
- https://www.together.ai/models/happyhorse-1-0-r2v
status: active
supportedModes:
- video
1 change: 1 addition & 0 deletions providers/together-ai/alibaba/happyhorse-1.0-t2v.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ modalities:
- text
output:
- video
- audio
mode: video
model: alibaba/happyhorse-1.0-t2v
provisioning: serverless
Expand Down
4 changes: 3 additions & 1 deletion providers/together-ai/alibaba/happyhorse-1.1-r2v.yaml
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
costs:
- input_cost_per_token: 0.0
- input_cost_per_request: 0.14
input_cost_per_token: 0.0
output_cost_per_token: 0.0
region: "*"
modalities:
Expand All @@ -16,6 +17,7 @@ removeParams:
- reasoning_effort
sources:
- https://www.together.ai/models/happyhorse-1-1-r2v
- https://www.together.ai/pricing
status: active
supportedModes:
- video
6 changes: 4 additions & 2 deletions providers/together-ai/alibaba/happyhorse-1.1-t2v.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,6 @@
costs:
- input_cost_per_request: 0.14
region: "*"
modalities:
input:
- text
Expand All @@ -9,8 +12,7 @@ model: alibaba/happyhorse-1.1-t2v
provisioning: serverless
sources:
- https://www.together.ai/models/happyhorse-1-1-t2v
- https://www.together.ai/pricing
status: active
supportedModes:
- video

# costs: pricing not found in official Together AI documentation (changelog, serverless models page, and pricing page only list happyhorse-1.0 variants)
1 change: 1 addition & 0 deletions providers/together-ai/black-forest-labs/FLUX-3.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ modalities:
- text
- image
- video
- audio
output:
- video
- audio
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,12 @@ modalities:
mode: image
model: black-forest-labs/FLUX.1-kontext-max
provisioning: serverless
removeParams:
- reasoning_effort
sources:
- https://docs.together.ai/docs/serverless-models
- https://docs.together.ai/docs/images-overview
- https://www.together.ai/models/flux-1-kontext-max
status: active
supportedModes:
- image
2 changes: 2 additions & 0 deletions providers/together-ai/deepcogito/cogito-v2-1-671b.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@ costs:
- input_cost_per_token: 0.00000125
output_cost_per_token: 0.00000125
region: "*"
features:
- function_calling
limits:
context_window: 163840
modalities:
Expand Down
14 changes: 13 additions & 1 deletion providers/together-ai/deepseek-ai/DeepSeek-V4-Flash-0731.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5,15 +5,27 @@ costs:
region: "*"
features:
- function_calling
- json_output
- structured_output
- tool_choice
- system_messages
limits:
context_window: 1048576
context_window: 1000000
modalities:
input:
- text
output:
- text
mode: chat
model: deepseek-ai/DeepSeek-V4-Flash-0731
params:
- defaultValue: high
key: reasoning_effort
supportedValues:
- low
- high
- max
type: string
provisioning: serverless
sources:
- https://www.together.ai/models/deepseek-v4-flash-0731
Expand Down
4 changes: 2 additions & 2 deletions providers/together-ai/google/flash-image-2.5.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -10,10 +10,10 @@ modalities:
mode: image
model: google/flash-image-2.5
provisioning: serverless
removeParams:
- reasoning_effort
sources:
- https://www.together.ai/pricing
- https://www.together.ai/models/gemini-flash-image-2-5
- https://docs.together.ai/docs/images-overview
status: active
supportedModes:
- image
3 changes: 3 additions & 0 deletions providers/together-ai/google/flash-image-3.1-lite.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,9 @@ costs:
- input_cost_per_token: 0.0 # not found in official docs
output_cost_per_token: 0.0 # not found in official docs
region: "*"
limits:
context_window: 1000000
max_input_tokens: 1000000
modalities:
input:
- text
Expand Down
4 changes: 3 additions & 1 deletion providers/together-ai/google/gemma-3-4b-pt.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
isDeprecated: true
limits:
context_window: 131072
max_output_tokens: 8192
Expand All @@ -13,9 +14,10 @@ model: google/gemma-3-4b-pt
provisioning: provisioned
removeParams:
- reasoning_effort
retirementDate: "2026-07-29"
sources:
- https://www.together.ai/models/gemma-3-4b
- https://huggingface.co/google/gemma-3-4b-pt
status: active
status: retired
supportedModes:
- completion
2 changes: 2 additions & 0 deletions providers/together-ai/google/imagen-4.0-fast.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ costs:
modalities:
input:
- text
- image
output:
- image
mode: image
Expand All @@ -12,6 +13,7 @@ provisioning: serverless
sources:
- https://www.together.ai/pricing
- https://docs.together.ai/docs/images-overview
- https://www.together.ai/models/google-imagen-4-0-fast
status: active
supportedModes:
- image
2 changes: 2 additions & 0 deletions providers/together-ai/google/veo-2.0.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,8 @@ modalities:
mode: video
model: google/veo-2.0
provisioning: serverless
removeParams:
- reasoning_effort
sources:
- https://www.together.ai/models/google-veo-2-0
status: active
Expand Down
5 changes: 2 additions & 3 deletions providers/together-ai/google/veo-3.1-lite.yaml
Original file line number Diff line number Diff line change
@@ -1,7 +1,6 @@
costs:
- input_cost_per_token: 0.0
output_cost_per_token: 0.0
region: "*" # pricing not listed on together.ai/pricing or together.ai/models for this model
- input_cost_per_request: 0.05
region: "*"
Comment thread
LordGameleo marked this conversation as resolved.
limits:
max_input_tokens: 1024
modalities:
Expand Down
10 changes: 8 additions & 2 deletions providers/together-ai/meta-llama/Meta-Llama-3.1-8B.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,7 @@
costs:
- input_cost_per_token: 1.8e-7
output_cost_per_token: 1.8e-7
region: "*"
limits:
context_window: 131072
max_input_tokens: 131072
Expand All @@ -11,9 +15,11 @@ modalities:
mode: completion
model: meta-llama/Meta-Llama-3.1-8B
provisioning: serverless
removeParams:
- reasoning_effort
sources:
- https://docs.together.ai/docs/serverless-models
- https://docs.together.ai/docs/deprecations
- https://www.together.ai/models/llama-3-1
- https://www.together.ai/pricing
Comment thread
LordGameleo marked this conversation as resolved.
status: active
supportedModes:
- completion
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ costs:
features:
- system_messages
- json_output
isDeprecated: true
limits:
context_window: 32768
modalities:
Expand All @@ -17,12 +18,13 @@ model: mistralai/Mistral-7B-Instruct-v0.2
provisioning: provisioned
removeParams:
- reasoning_effort
retirementDate: "2026-07-29"
sources:
- https://www.together.ai/models/mistral-7b-instruct-v0-2
- https://www.together.ai/pricing
- https://docs.together.ai/docs/function-calling
- https://docs.together.ai/docs/json-mode
- https://docs.together.ai/docs/dedicated-models
status: active
status: retired
supportedModes:
- chat
1 change: 1 addition & 0 deletions providers/together-ai/moonshotai/Kimi-K2.7-Code.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ features:
- tool_choice
- system_messages
- structured_output
- prompt_caching
limits:
context_window: 262144
modalities:
Expand Down
2 changes: 1 addition & 1 deletion providers/together-ai/moonshotai/Kimi-K3.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ features:
- structured_output
- prompt_caching
limits:
context_window: 1000000
context_window: 1048576
modalities:
input:
- text
Expand Down
2 changes: 2 additions & 0 deletions providers/together-ai/nvidia/nemotron-3-ultra-550b-a55b.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,8 @@ modalities:
mode: chat
model: nvidia/nemotron-3-ultra-550b-a55b
provisioning: serverless
removeParams:
- reasoning_effort
sources:
- https://www.together.ai/models/nvidia-nemotron-3-ultra
status: active
Expand Down
1 change: 1 addition & 0 deletions providers/together-ai/pearl-ai/gemma-4-31b-it.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ costs:
features:
- function_calling
- json_output
- tool_choice
limits:
context_window: 32000
modalities:
Expand Down
2 changes: 2 additions & 0 deletions providers/together-ai/prunaai/p-image-ideogram.yaml
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
costs:
- input_cost_per_token: 0.0
output_cost_per_image: 0.00225
output_cost_per_token: 0.0
region: "*"
modalities:
Expand All @@ -12,5 +13,6 @@ model: prunaai/p-image-ideogram
provisioning: serverless
removeParams:
- reasoning_effort
status: active
supportedModes:
- image
14 changes: 13 additions & 1 deletion providers/together-ai/thinkingmachines/Inkling-Small.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -3,8 +3,10 @@ costs:
input_cost_per_token: 5e-7
output_cost_per_token: 1.2e-6
region: "*"
features:
- function_calling
limits:
context_window: 524288
context_window: 1000000
modalities:
input:
- text
Expand All @@ -14,6 +16,16 @@ modalities:
- text
mode: chat
model: thinkingmachines/Inkling-Small
params:
- defaultValue: medium
key: reasoning_effort
supportedValues:
- minimal
- low
- medium
- high
- xhigh
type: string
Comment thread
LordGameleo marked this conversation as resolved.
provisioning: serverless
sources:
- https://www.together.ai/models/inkling-small
Expand Down
Loading