Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
search_modelsA

Unified endpoint for discovering model endpoints. Supports three usage modes:

1. List Mode (no parameters): Paginated list of all available model endpoints with minimal metadata.

2. Find Mode (endpoint_id parameter): Retrieve specific model endpoint(s) by ID. Supports single or multiple IDs.

3. Search Mode (search parameters): Filter models by free-text query, category, or status.

Expansion: Use expand to include additional data in each model object:

  • openapi-3.0 — full OpenAPI 3.0 schema in the openapi field

  • enterprise_status — enterprise readiness status (ready or pending) in the enterprise_status field

Examples of endpoint_id values:

  • fal-ai/flux/dev

  • fal-ai/wan/v2.2-a14b/text-to-video

  • fal-ai/minimax/video-01/image-to-video

  • fal-ai/hunyuan3d-v21

See fal.ai Model APIs for more details.

Authentication: Optional. Providing an API key grants higher rate limits.

Common Use Cases:

  • Browse available models for integration

  • Retrieve metadata for specific endpoints

  • Search for models by category or keywords

  • Get OpenAPI schemas for code generation

  • Build model selection interfaces

get_pricingA

Returns unit pricing for requested endpoint IDs. Most models use output-based pricing (e.g., per image/video with proportional adjustments for resolution/length). Some models use GPU-based pricing depending on architecture. Values are expressed per model's billing unit in a given currency.

Authentication: Required. Users must provide a valid API key. Custom pricing or discounts may be applied based on account status.

Common Use Cases:

  • Display pricing in user interfaces

  • Compare pricing across different models

  • Build cost estimation tools

  • Check current billing rates

See fal.ai pricing for more details.

estimate_pricingA

Computes cost estimates using one of two methods:

1. Historical API Price (historical_api_price):

  • Based on historical pricing per API call from past usage patterns

  • Takes call_quantity (number of API calls) per endpoint

  • Useful for estimating based on actual historical usage patterns

  • Example: "How much will 100 calls to flux/dev cost?"

2. Unit Price (unit_price):

  • Based on unit price × expected billing units from pricing service

  • Takes unit_quantity (number of billing units like images/videos) per endpoint

  • Useful when you know the expected output quantity

  • Example: "How much will 50 images from flux/dev cost?"

Authentication: Required. Users must provide a valid API key. Custom pricing or discounts may be applied based on account status.

Common Use Cases:

  • Pre-calculate costs for batch operations

  • Display cost estimates in user interfaces

  • Budget planning and cost optimization

See fal.ai pricing for more details.

get_usageA

Returns paginated usage records for your workspace with filters for endpoint, user, date range, and auth method. Each item includes the billed unit quantity, the pre-discount unit price and cost_subtotal, any percentage discount applied, and the final cost_total (cost_subtotal − cost_discount).

Key Features:

  • Usage data for all endpoints or filtered by specific endpoint(s)

  • Flexible date range filtering

  • User-specific usage tracking

  • Detailed usage line items with unit quantity, price, and discount breakdown

  • Paginated results for large datasets

Common Use Cases:

  • Generate usage reports for all endpoints or specific models

  • Track usage patterns

  • Monitor endpoint usage across different auth methods

  • Build usage dashboards and visualizations

See fal.ai docs for more details.

get_analyticsA

Time-bucketed metrics per model endpoint, including request counts, success/error rates, and latency percentiles. prepare_duration reflects queue/prepare time before execution; duration is request execution time. Use with the Queue/Webhooks flow to monitor SLAs.

Metric Selection: You must specify which metrics to include using the expand query parameter. Only requested metrics will be populated in the response, allowing you to optimize query performance and data transfer.

Available Metrics:

The expand parameter accepts these values, grouped by category:

Volume

  • request_count: Total number of requests in the time bucket

  • success_count: Successful requests (2xx responses)

  • user_error_count: User errors (4xx responses)

  • error_count: Server errors (5xx responses)

Error type breakdown

  • startup_error_count: Startup errors (startup timeout, scheduling failure)

  • connection_error_count: Connection errors (timeout, disconnected, refused)

  • timeout_error_count: Request timeout errors

  • runtime_error_count: Runtime errors (internal error, server error)

Queue / prepare latency

  • p50_prepare_duration, p75_prepare_duration, p90_prepare_duration, p95_prepare_duration, p99_prepare_duration: Time from request submission until execution starts

Request execution latency

  • p25_duration, p50_duration, p75_duration, p90_duration, p95_duration, p99_duration: Time spent processing the request

Cold boot

  • cold_boot_count: Requests with cold boot (startup > 1s)

  • p50_cold_boot_duration, p75_cold_boot_duration, p90_cold_boot_duration: Cold boot duration percentiles

Billing

  • total_billable_duration: Aggregate billed execution time

Key Features:

  • Selective metric inclusion via expand parameter

  • Performance metrics (latency percentiles, duration stats)

  • Reliability metrics (success/error rates, request counts)

  • Error type breakdown (startup, connection, timeout, runtime)

  • Cold boot metrics (count, latency percentiles)

  • Billing duration tracking

  • Time-bucketed data for trend analysis

  • Single or multi-model analytics

  • Flexible date range and timeframe options

Common Use Cases:

  • Monitor model performance and reliability

  • Generate performance dashboards

  • Analyze latency trends and patterns

  • Track error rates and success metrics

See Queue API docs for more details.

get_billing_eventsA

Returns paginated individual billing event records with filters for endpoint and date range. Each record includes the request ID, timestamp, endpoint, output units billed, and a cost breakdown in USD (cost_subtotal, cost_discount, cost_total; cost_estimate_nano_usd carries cost_total in nano USD).

Key Features:

  • Individual billing event records for each API request

  • Per-request cost breakdown before and after discounts

  • Flexible date range filtering

  • Optional endpoint filtering

  • Cursor-based pagination for efficient large dataset queries

  • Limited to 10000 records per page for performance

  • Date range capped at 90 days per request

Common Use Cases:

  • Audit individual billing events

  • Track request patterns and volumes

  • Debug specific requests by ID

  • Monitor billing unit consumption per request

See fal.ai docs for more details.

delete_request_payloadsA

Deletes the IO payloads and associated CDN output files for a specific request.

Important:

  • Only output CDN files are deleted (input files may be used by other requests)

  • This action is irreversible

  • Requires authentication with an admin API key

What gets deleted:

  • Request input/output payload data

  • CDN-hosted output files (images, videos, etc.)

What is NOT deleted:

  • Input CDN files (may be referenced by other requests)

Response:

  • Returns deletion status for each CDN file

  • Each result includes the file link and any error that occurred

Idempotency:

  • Optional Idempotency-Key header prevents duplicate deletions on retries

  • Responses cached for 10 minutes per unique key

See fal.ai docs for more details about request payloads.

list_requests_by_endpointA

Lists requests for one or more endpoints (same endpoint_id style as usage/explore: comma-separated or repeated query params, up to 50 IDs).

Authentication: Requires API key (user or enterprise).

Filters:

  • Time range via start / end. If start is omitted, defaults to the last 24 hours — unless request_id is provided, in which case the default start bound is widened to 90 days.

  • Status (success, error, user_error)

  • Request ID

  • Pagination via cursor/limit (limit defaults to 50, max 100)

Sorting:

  • By end time (default) or duration

Expansions:

  • Include payloads by adding expand=payloads

search_requestsA

Search, filter, and browse your request history. Supports three modes:

1. Semantic Search (query, image_url, or video_url parameter): Find visually or conceptually similar results using AI embeddings. Provide a text query for text-to-image search, an image URL for image-to-image similarity search, or a video URL for video-to-image similarity search.

2. Filtered Browse (no query, image_url, or video_url): Browse request history with hard filters. Returns results ordered by creation date (newest first).

3. Semantic + Filters (search params AND filter params): Combine semantic search with hard filters. Filters narrow the candidate set before ranking by similarity.

Filter Options:

  • endpoint_id: Filter by one or more fal endpoints (comma-separated or repeated, up to 50 IDs)

  • exclude_api_requests / only_api_requests: Filter by request source

Examples:

  • Semantic text search: ?query=sunset+landscape

  • Image similarity: ?image_url=https://...&min_similarity=0.5

  • Filtered search: ?query=portrait&endpoint_id=fal-ai/flux/dev

  • Browse across multiple endpoints: ?endpoint_id=fal-ai/flux/dev,fal-ai/flux/schnell

list_workflowsA

List workflows for the authenticated user with optional search and filtering.

Features:

  • Paginated results with cursor-based pagination

  • Search by workflow name or title

  • Filter by model endpoints used in the workflow

Authentication: Required. Returns only workflows owned by the authenticated user.

Common Use Cases:

  • Display user's workflow library

  • Search for specific workflows

  • Find workflows using particular models

create_workflowA

Create a new workflow owned by the authenticated user.

Authentication: Required.

Common Use Cases:

  • Save a newly built workflow

  • Programmatically provision workflows

Note: Workflow names must be unique within your namespace. Creating a workflow with a name you already use returns a 400 validation error.

get_workflowA

Get detailed information about a specific workflow, including its full contents/definition.

Authentication: Required.

Common Use Cases:

  • Load a workflow for editing

  • View workflow configuration

  • Export workflow definition

list_assetsB

Browse and semantically search fal Assets across all media, uploads, favorites, collections, tags, and character references.

list_asset_collectionsB

List asset collections for the authenticated user's fal Assets library.

create_asset_collectionC

Create asset collection for the authenticated user's fal Assets library.

get_asset_collectionC

Get asset collection for the authenticated user's fal Assets library.

update_asset_collectionC

Update asset collection for the authenticated user's fal Assets library.

delete_asset_collectionB

Delete asset collection for the authenticated user's fal Assets library.

get_asset_collection_hierarchyA

Get the nested subtree rooted at an asset collection, plus its ancestor collections ordered from the top level down to its direct parent.

favorite_asset_collectionB

Favorite an asset collection for the authenticated user's fal Assets library.

unfavorite_asset_collectionB

Unfavorite an asset collection for the authenticated user's fal Assets library.

move_asset_collectionA

Move a manual asset collection under another collection, or to the top level. Only manual collections can be moved or act as folders; nesting is limited to 5 levels deep and cannot create a cycle.

list_asset_collection_assetsC

Browse assets in a collection for the authenticated user's fal Assets library.

add_asset_to_collectionA

Add an asset to a manual or character collection. Provide a request ID or vector ID; unresolved references are materialized before local collection state is added. For character collections, the asset is added by applying the character tag.

remove_asset_from_collectionA

Remove an asset from a manual or character collection by request ID or vector ID.

list_asset_charactersB

List asset characters for the authenticated user's fal Assets library.

create_asset_characterB

Create an asset character for the authenticated user's fal Assets library. Prefer vector IDs or request IDs in reference_images for existing fal-generated assets; use fal-hosted image URLs only for standalone images. Unresolved ID references are materialized before character state is added.

update_asset_characterA

Update an asset character for the authenticated user's fal Assets library. Prefer vector IDs or request IDs in reference_images for existing fal-generated assets; use fal-hosted image URLs only for standalone images. Unresolved ID references are materialized before character state is added.

get_asset_characterC

Get asset character for the authenticated user's fal Assets library.

delete_asset_characterB

Delete asset character for the authenticated user's fal Assets library.

favorite_asset_characterC

Favorite an asset character for the authenticated user's fal Assets library.

unfavorite_asset_characterB

Unfavorite an asset character for the authenticated user's fal Assets library.

list_asset_tagsB

List asset tags for the authenticated user's fal Assets library.

create_asset_tagC

Create asset tag for the authenticated user's fal Assets library.

set_asset_tags_for_assetB

Set tags for an asset. Provide a request ID or vector ID; unresolved references are materialized before tag state is added.

update_asset_tagC

Update asset tag for the authenticated user's fal Assets library.

delete_asset_tagB

Delete asset tag for the authenticated user's fal Assets library.

upload_assetB

Upload asset for the authenticated user's fal Assets library.

get_assetA

Get an asset document by vector ID from the authenticated user's fal Assets library. The vector may exist only in Turbopuffer; in that case the response returns the Turbopuffer document with empty local state.

get_asset_lineageA

Get the derivation lineage of an asset by asset ID: the inputs it was generated from, the generation requests along the way, and any referenced characters, traversed recursively up to depth levels. Deleted or expired ancestors stay in the graph flagged as tombstones; inputs that were never captured appear as external inputs.

favorite_assetA

Favorite an asset. Provide a request ID or vector ID; unresolved references are materialized before favorite state is added.

unfavorite_assetC

Unfavorite an asset by request ID or vector ID.

list_asset_tags_for_assetA

List tags for an asset by vector ID. Vectors that have not been saved as assets return an empty tag list.

assign_asset_tagA

Assign a tag to an asset. Provide a request ID or vector ID; unresolved references are materialized before tag state is added.

unassign_asset_tagC

Unassign a tag from an asset by request ID or vector ID.

get_storage_file_aclA

Returns the Access Control List currently applied to a fal CDN file.

The ACL consists of a default decision (allow, forbid, or hide) plus optional per-user rules that override the default. Rule users are returned as nicknames where possible.

Authentication: Required. The API key must have the assets:read permission.

set_storage_file_aclA

Replaces the Access Control List of a fal CDN file.

The ACL consists of a default decision (allow, forbid, or hide) plus optional per-user rules that override the default. Rule users may be specified by nickname or user ID. Setting default to allow with no rules makes the file public; forbid or hide restricts it to the rules you provide.

Rules referencing users that do not exist are dropped. The response reflects the ACL actually applied, so verify it contains the rules you sent.

Authentication: Required. The API key must have the assets:write permission.

sign_storage_file_urlA

Creates a signed URL that grants temporary access to a fal CDN file, regardless of its ACL. Useful for sharing access-restricted files.

The signature is valid for expiration_seconds (up to 7 days).

Authentication: Required. The API key must have the assets:read permission.

get_storage_settingsA

Returns the account-level storage lifecycle settings applied to newly uploaded fal CDN files:

  • expiration_duration_seconds: how long files live before being automatically deleted (null disables auto-expiration).

  • initial_acl: the default ACL applied to new uploads (null means the system default, which is public).

Both fields are null when the account has never saved settings.

Authentication: Required. The API key must have the account:settings:read permission.

update_storage_settingsA

Replaces the account-level storage lifecycle settings applied to newly uploaded fal CDN files. Omitted or null fields are cleared (reset to the system default), so always send the full desired configuration.

ACL rules referencing users that do not exist are dropped. The response reflects the settings actually saved, so verify it contains the rules you sent.

These are the same settings that the per-request X-Fal-Object-Lifecycle-Preference header overrides on individual requests.

Authentication: Required. The API key must have the account:settings:write permission.

get_account_billingB

Returns billing information for the authenticated account. Use the expand parameter to include additional details.

Expandable Fields:

  • credits — Current credit balance and currency

Common Use Cases:

  • Monitor available credit balance programmatically

  • Display balance in custom dashboards

get_organization_teamsA

Returns the list of teams in your organization with their details.

Availability: This endpoint is available to enterprise customers with organizations enabled. Contact your account team or support@fal.ai to request access.

Must be called with an admin API key on the organization's root team.

Key Features:

  • List all teams within the organization

  • Identify the organization's root team via is_org_root

  • View team usernames and display names

See fal.ai docs for more details.

get_organization_usageA

Returns paginated usage records across all teams and product lines in your organization, with each record attributed to a specific team via the username field and a product line via the product field.

Covers all three fal product lines:

  • model_apis — model API endpoint calls (e.g. fal-ai/flux/dev)

  • serverless — fal Serverless SDK billing

  • compute — fal Compute (raw instance time)

Availability: This endpoint is available to enterprise customers with organizations enabled. Contact your account team or support@fal.ai to request access.

Must be called with an admin API key on the organization's root team.

Key Features:

  • Organization-wide usage data across all teams and products

  • Filter by team(s) (team_username), product line (product), endpoint, API key (api_key_id), date range, and auth method

  • Per-team and per-product attribution on every usage record

  • Paginated time series and aggregate summary views

See fal.ai docs for more details.

get_model_infoC

Exact current catalog lookup with OpenAPI expansion. No generation or inferred model defaults. Schema may be unavailable; inspect actual native fields.

run_modelA

Confirmed paid model request after current native input-schema validation. Exact model_id/input required; image/video commands do not invent fields or choose a default. Queue submissions return receipt only; synchronous timeout may leave an unknown paid outcome. No retry or polling.

submit_jobA

Confirmed paid model request after current native input-schema validation. Exact model_id/input required; image/video commands do not invent fields or choose a default. Queue submissions return receipt only; synchronous timeout may leave an unknown paid outcome. No retry or polling.

generate_imageB

Confirmed paid model request after current native input-schema validation. Exact model_id/input required; image/video commands do not invent fields or choose a default. Queue submissions return receipt only; synchronous timeout may leave an unknown paid outcome. No retry or polling.

generate_videoB

Confirmed paid model request after current native input-schema validation. Exact model_id/input required; image/video commands do not invent fields or choose a default. Queue submissions return receipt only; synchronous timeout may leave an unknown paid outcome. No retry or polling.

get_job_statusB

One read using the SDK-compatible owner/app root, not the full model subpath. No auto-polling, paid re-submission or arbitrary status URL.

get_job_resultA

One result read using the receipt/model root. No wait loop, media download, re-submission or auto-upload. Signed credential URLs are redacted; ordinary output media URLs remain account data.

cancel_jobB

Confirmed native cancellation request. Cancellation receipt is not proof processing stopped or credits were refunded; current provider state controls eligibility. No retry.

upload_fileB

Confirmed selected absolute regular non-symlink local file, 1 byte–20 MiB. Uses pinned SDK upload-initiation protocol then a credential-free HTTPS fal.media PUT with redirects refused. No remote URL ingestion, base64 model output, multipart retries or automatic generation.

list_accountsB

Local labels/default/auth method only. No keys, token paths, real provider identities or network request.

get_operation_schemaC

Local current native method/path/query/header/body schema and exact provenance. No credential or provider request.

preview_generation_batchA

Read-only current schema validation and native unit-pricing lookup for all requested async jobs. Hash binds ordered exact inputs/lifecycle/store-IO/profile label/current schemas/unit quotes. Unit pricing is not final cost or a spending cap. No generation, file write or key ownership validation.

submit_generation_batchA

Confirmed one-to-ten async jobs. Refetch all current schemas/unit quotes and validate all before first paid submission; refuse changed hash. Submit sequentially, stop on first failure, report known request IDs/failed and unattempted indices. No polling, retries, rollback, continuation or budget guarantee.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.1/5.0

Scored across 66 tools

Disambiguation2/5

Several tools have overlapping or identical purposes: run_model, submit_job, generate_image, and generate_video all share the exact same description, making it impossible to distinguish when to use which. search_requests and list_requests_by_endpoint also overlap heavily, and list_assets vs list_asset_collection_assets vs get_asset cover similar ground.

Naming Consistency4/5

Most tools follow a consistent verb_noun snake_case pattern (get_model_info, create_workflow, delete_asset_collection). A few deviations like assign_asset_tag/unassign_asset_tag vs set_asset_tags_for_asset add minor inconsistency but the overall scheme is predictable.

Tool Count2/5

66 tools is far too many for coherent agent use; the set is bloated with many near-duplicate CRUD operations for assets, collections, characters, and tags. This volume forces significant disambiguation burden on the caller.

Completeness3/5

The surface covers many domains (models, pricing, usage, analytics, billing, workflows, assets, storage, jobs) with reasonable CRUD completeness for assets and collections. However, clear gaps exist: no list_workflows deletion/update, no search_assets tool despite list_assets mentioning semantic search, and no clear way to poll/cancel batch jobs beyond the single cancel_job.

Maintenance

ActivityNo data
ResponsivenessNo issues