FLA Dispatch Backends Skill
SkillProductivityGuides your agent through writing and updating backend dispatch code in the flash-linear-attention library.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the FLA Dispatch Backends Skill skill
About this capability
Workflow for FLA backend dispatch decorators and backend implementations. Use when touching fla.ops.backends, @dispatch-decorated functions, BaseBackend subclasses, backend verifier methods, backend env vars, or backend tests.
What this skill tells your AI
The instructions your AI receives, as published by fla-org/flash-linear-attention in .agents/skills/fla-dispatch-backends/SKILL.md and read by ahel’s review.
Use this skill for the runtime backend dispatch system implemented in
fla/ops/backends/__init__.py.
Core model
- Public functions opt in with
@dispatch('<operation>'). - First call lazily imports
fla.ops.<operation>.backends, unless the operation has a custom module in_OPERATION_BACKEND_MODULES(for examplemodules). - Backend modules create
BackendRegistry('<operation>')and registerBaseBackendsubclasses. - Dispatch tries registered backends sorted by
prioritywhere lower means higher priority. - A backend is considered only when
is_available()andis_enabled()are both true. - Runtime dispatch checks
is_available()andis_enabled()directly; do not rely on the cachedcan_use()path inside code that must be torch.compile friendly. - If
<func_name>_verifierexists, it must return(True, None)or(False, reason). Rejected calls fall back to the next backend. - If no backend handles the call, dispatch runs the original implementation.
FLA_DISABLE_BACKEND_DISPATCH=1bypasses the decorator entirely.- The dispatch wrapper is marked with
torch.compiler.disable, so keep backend selection logic outside compiled graphs and keep compiled work inside the selected backend implementation.
Backend implementation checklist
For a new backend:
- Add a
BaseBackendsubclass under the operation'sbackends/package. - Set
backend_type,package_name,env_var,default_enable, andpriority. - Implement
<public_function_name>_verifier(...)with the same public call surface as the decorated function. - Implement
<public_function_name>(...)and keep return values identical to the default implementation. - Register the backend in the operation's
backends/__init__.py. - Add tests that cover accepted dispatch, verifier rejection, and fallback.
Verifier rules
- Verifiers must be cheap, deterministic, and side-effect free.
- Return a specific rejection reason; it is logged once and is useful in CI logs.
- Check dtype, shape, layout, inference/training mode, external package requirements, env flags, and unsupported options before calling backend code.
- Do not silently copy or normalize inputs in a verifier; do that in the backend implementation only when it is part of the backend contract.
- If a backend supports only inference, check
torch.is_grad_enabled()ortorch.is_inference_mode_enabled()as appropriate. - Do not mutate global backend registries, environment variables, tensors, RNG state, or caches from a verifier.
Decorator placement
- Decorate public operation entry points, not private helpers that are only used inside one backend.
- Keep the decorated function as the semantic fallback implementation. A user
should be able to set
FLA_DISABLE_BACKEND_DISPATCH=1and still get the same API behavior. - Use the operation name that maps to the backend package. For normal ops,
@dispatch('kda')maps tofla.ops.kda.backends; special cases belong in_OPERATION_BACKEND_MODULES. - Do not add import-time side effects in backend packages beyond registering backends.
Testing guidance
- Test the public decorated function, not only the backend helper.
- Force dispatch off with
FLA_DISABLE_BACKEND_DISPATCH=1when comparing against the Triton/default path. - Force or disable backend-specific env vars (
FLA_FLASH_KDA,FLA_TILELANG,FLA_INTRACARD_CP) when testing route behavior. - Include at least one rejection test for each verifier branch added or changed.
- For backend changes under
fla/ops/<op>/backends/, ensure dependent op tests still run;scripts/find_dependent_tests.pymaps backend changes back to the decorated op files.
Style constraints
- Use platform helpers from
fla.utilsfor hardware/platform decisions instead of adding new directtorch.cudachecks in public code or tests. If no helper covers the condition, add a small helper infla.utilsfirst. - Keep backend imports lazy inside backend implementations when importing an optional package would otherwise break environments without that package.
- Keep error/rejection messages precise and user-facing; they appear in logs and tests may assert them.
Signals
- GitHub stars
- 6k
- Forks
- 702
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
fla-dispatch-backends- Source
- github.com/fla-org/flash-linear-attention