fal (Flux image generation)
fal.ai serves the Flux image models. Unlike the chat providers, image generation is billed per image rather than per token — which makes it very easy to lose track of, because a single workflow run can produce several images and none of it shows up in a token count.
Connect fal to TokenSense
Step 1: Go to your dashboard → Settings → Providers → add your fal.ai API key.
Step 2: Call /v1/images/generations on your TokenSense endpoint with a flux- model name.
Step 3: Authenticate with your TokenSense API key. Each generated image is logged with its cost and attributed to the workflow that requested it.
Supported models
Pricing is per image (USD).
| Model | Cost per image | Notes |
|---|---|---|
| flux-2-pro | $0.050 | Highest quality |
| flux-2-dev | $0.025 | Balanced |
| flux-2-schnell | $0.003 | Fastest and cheapest |
Per-image costs add up faster than they look. A workflow generating 4 images per run on flux-2-pro costs $0.20 a run — at 200 runs a day that is $40 a day, roughly $1,200 a month from one workflow. The same workflow on flux-2-schnell is about $72 a month. Worth checking which model a workflow actually needs before it scales.
What fal does not support
fal is image generation only. It has no chat completions, embeddings, or audio models, so /v1/chat/completions, /v1/embeddings, and the audio endpoints need a different provider. See the provider overview for what covers each modality.
Using fal in your workflows
n8n: Use an HTTP Request node pointed at /v1/images/generations on your TokenSense endpoint. See the n8n setup guide.
Make / Zapier: Use an HTTP module against the same endpoint. See the Make & Zapier guide.
Custom code: Point your OpenAI client's base URL at your TokenSense endpoint and call the images endpoint with a flux- model. See the API reference.
Related
- OpenAI Provider Guide — DALL-E and gpt-image-1
- Google Gemini Provider Guide — image adapter and lifecycle guidance
- Budgets & Spend Caps
- Cost Attribution
