Production FastAPI Template: Rules
SkillCloud & infraStructure and execution rules for a unified, production-ready backend combining domain-oriented FastAPI, Retrieval-Augmented Generation (RAG) and Agentic AI in one codebase on AWS ECS Fargate. Use when scaffolding a new backend; adding a module, API domain, RAG component, agent, tool, workflow, LLM/provider integration or test; deciding where a file belongs; writing async/sync routes, services or clients; offloading blocking or CPU-heavy work (thread pools, process pools, AWS Lambda); designing background jobs, queues (SQS, Celery), workers and scaling; or writing Pydantic schemas, settings (pydantic-settings), serialization, structured LLM outputs, citation validation, agent plans or tool-call arguments; or writing FastAPI dependencies (request validation, auth, access scopes, chaining), REST paths, response models, OpenAPI docs, error responses, database naming and SQL queries, migrations, API tests (async client, dependency overrides) or lint/format setup (ruff); or adding content-safety guardrails (input/output checks, PII, prompt injection, toxicity, Bedrock Guardrails, AgentCore Policy/Gateway Cedar policies, Guardrails AI validators), fail-closed handling and guardrail evaluation; or evaluating RAG, retrieval stages, structured outputs, citations, agents, multi-agent workflows or conversations (DeepEval, Ragas, LLM-as-judge, G-Eval, golden datasets, regression tiers, thresholds, judge calibration); or securing the API (authentication, JWT/JWKS, OAuth 2.0/OIDC, RBAC/ABAC, BOLA/BFLA, tenant isolation, input/upload/SSRF validation, security headers, CORS, CSRF, rate limiting, audit logging, ECS/IAM/secrets hardening, CI security scanning, security tests) or LLM/agent security (OWASP LLM Top 10, OWASP Agentic Top 10, prompt injection, tool authorization, human approval, excessive agency).
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Production FastAPI Template: Rules skill
What this skill tells your AI
The instructions your AI receives, as published by davepoon/buildwithclaude in plugins/production-fast-api-template/skills/production-fast-api-template/SKILL.md and read by ahel’s review.
These rules define where code lives and how it executes in a unified FastAPI + RAG + Agentic AI backend. Apply them whenever you generate or modify code in a project built from this template.
Companion references. Read the relevant file before writing code:
| File | Read before writing… | Summarized in |
|---|---|---|
| ASYNC_EXECUTION.md | Any route, service, client, agent tool, worker or background job | Section 5 |
| PYDANTIC_STANDARDS.md | Any schema, settings class, serialization, LLM call with structured output, citation check, agent plan or tool argument model | Section 6 |
| API_CONVENTIONS.md | Any dependency, route signature, auth/permission check, OpenAPI metadata, DB model or query, migration, API test or lint setup | Section 7 |
| GUARDRAILS.md | Any user-facing LLM call, RAG answer, ingestion of untrusted documents, agent tool execution, gateway target or content-safety check | Section 8 |
| AI_EVALUATION.md | Any evaluation contract, metric, dataset, judge, runner, report or threshold; any new RAG stage, agent, tool or conversation feature that needs quality coverage | Section 9 |
| API_SECURITY.md | Any authentication, token, permission, tenant or ownership check; input/upload/outbound-URL handling; response headers, CORS, errors; rate limits; audit events; ECS/IAM/secrets; CI security; security tests | Section 10 |
| AGENT_SECURITY.md | Any agent, tool, tool registry entry, approval flow, memory write, MCP connector or LLM-output sink | Section 10 |
Status labels:
- 🟢 Convention. A proposed rule. Follow it by default.
- 🟡 Needs approval. An open team decision. Surface the options to the user and never pick one silently.
These rules don't choose agent frameworks, LLM providers, models, vector stores or a queue implementation. If a rule doesn't fit a real need, raise it with the user instead of silently working around it.
1. Core structure rules 🟢
- One application. FastAPI, RAG and Agentic AI are not separate apps. There is exactly one
src/package, one FastAPI entry point (src/main.py) and one shared infrastructure layer. - Never create a second entry point. Don't add a second
src/, a separate app folder (rag_app/,agents_service/,worker_app/) or an extraFastAPI()instance. - Background workers and Lambda functions are extra entry points into the same codebase. They are not separate applications, and they call the same service functions the API calls.
- One central API router. All routers mount through
src/api/v1/router.py. Domain, RAG and agent routers never mount themselves on the app. - Domain-oriented modules. Business code lives inside its domain package.
src/'s root holds only application-wide concerns. - No duplicated shared layers. Model clients go in
src/llm/, content-safety checks insrc/guardrails/and external-service clients insrc/providers/, each exactly once. RAG and agents consume them and never define their own copies (norag/llm_client.py, noagents/vector_db.py). - Provider-neutral naming. Name files by responsibility, not vendor (
llm/embeddings.py, notopenai_embeddings.py). Vendor specifics stay behindsrc/providers/orsrc/llm/. common/is for genuinely reusable, domain-agnostic code only. Never put business logic there.- Keep agents, reasoning and execution apart. Agent definitions go in
agents/, thinking incognition/and doing, scheduling and offloading inexecution/. - Tests mirror
src/. Quality evaluation (datasets, DeepEval/Ragas suites, deterministic metrics, runners, reports) lives in the top-levelevaluation/directory, separate fromtests/.src/evaluation/holds only framework-neutral contracts, the metric registry and adapters, with lazy framework imports (AI_EVALUATION.md §1). - Every Python directory under
src/has an__init__.py. Non-Python directories that must survive in Git (prompt folders, data, empty test folders) get a.gitkeep. - Explicit module imports across packages, e.g.
from src.auth import constants as auth_constants. - Reuse existing files before adding new ones. Extend
execution/background_worker.pyrather than creatingexecution/worker2.pyorrag/worker.py.
2. Canonical structure
<project-name>/
├── src/
│ ├── __init__.py
│ ├── main.py # the single FastAPI entry point (lifespan creates shared clients/limiters)
│ ├── config.py # application-wide configuration
│ ├── constants.py
│ ├── exceptions.py # base domain exception + handlers rendering ErrorResponse
│ ├── middleware.py
│ ├── database.py # shared metadata (naming convention), engine/session factory, get_db_session
│ ├── models.py # genuinely shared models only
│ ├── pagination.py
│ │
│ ├── api/
│ │ └── v1/
│ │ └── router.py # aggregates domain, RAG and agent routers
│ │
│ ├── auth/ # reference domain module
│ │ ├── router.py
│ │ ├── schemas.py # the domain's own Pydantic API/internal schemas
│ │ ├── models.py # persistence (ORM) models, not Pydantic
│ │ ├── dependencies.py # current_principal, require_<permission>, scope dependencies
│ │ ├── config.py # the domain's own pydantic-settings class (AUTH_ prefix: issuers, audiences, JWKS)
│ │ ├── constants.py # permission and scope names
│ │ ├── exceptions.py
│ │ ├── service.py
│ │ ├── tokens.py # JWT/JWKS validation, token types, claims → Principal
│ │ ├── permissions.py # RBAC/ABAC policy evaluation (default deny)
│ │ ├── security.py # password hashing, only if local accounts are approved 🟡
│ │ ├── audit.py # auth audit events (login, token rejected, session revoked)
│ │ └── utils.py
│ │
│ ├── example_domain/ # copy this layout for every new business domain
│ │ ├── router.py
│ │ ├── schemas.py
│ │ ├── models.py
│ │ ├── dependencies.py # valid_<entity>_id and other I/O-backed validation
│ │ ├── constants.py
│ │ ├── exceptions.py
│ │ ├── service.py
│ │ └── utils.py
│ │
│ ├── rag/
│ │ ├── routes.py # "routes", not "router": avoids clashing with rag/router/
│ │ ├── schemas.py # RAG API schemas + cross-stage contracts (RetrievedEvidence)
│ │ ├── service.py # container/lifecycle of long-lived RAG services
│ │ ├── dependencies.py
│ │ ├── config.py # RAG settings (RAG_ prefix)
│ │ ├── constants.py
│ │ ├── exceptions.py
│ │ ├── extraction.py # structure-preserving document extraction
│ │ ├── ingestion.py # coordinates processing and indexing (run by background jobs)
│ │ ├── semantic/
│ │ │ ├── metadata.py
│ │ │ ├── segmentation.py
│ │ │ └── chunks.py
│ │ ├── router/
│ │ │ ├── classifier.py
│ │ │ ├── fusion.py
│ │ │ └── evaluation.py
│ │ ├── preprocessing/
│ │ │ ├── query_expansion.py
│ │ │ └── hyde.py
│ │ ├── retrieval/
│ │ │ ├── semantic_search.py
│ │ │ ├── bm25.py
│ │ │ ├── rrf.py
│ │ │ ├── mmr.py
│ │ │ ├── evidence_filter.py
│ │ │ └── hybrid_retriever.py
│ │ ├── generation/
│ │ │ ├── context_builder.py
│ │ │ ├── citations.py
│ │ │ ├── streaming.py
│ │ │ ├── schemas.py # LLM-output contracts: GeneratedAnswer, Citation
│ │ │ ├── validation.py # application-level answer/citation validation
│ │ │ └── service.py
│ │ ├── chat/
│ │ │ ├── routes.py
│ │ │ ├── schemas.py
│ │ │ ├── service.py
│ │ │ └── storage.py
│ │ └── security/
│ │ ├── access_control.py # AccessScope model + vector-store/DB filter translation
│ │ ├── document_validation.py # ingestion file checks (type, signature, size, archives)
│ │ └── retrieval_policy.py # post-retrieval ACL re-check, trust labels, provenance
│ │
│ ├── agents/
│ │ ├── router.py
│ │ ├── schemas.py # AgentTask, ToolResult, AgentExecutionState
│ │ ├── structured_output.py # LLM-produced: ExecutionPlan, ToolCall union, FinalAgentResponse
│ │ ├── dependencies.py
│ │ ├── config.py # agent settings (AGENTS_ prefix)
│ │ ├── constants.py
│ │ ├── exceptions.py
│ │ ├── service.py
│ │ ├── base_agent.py
│ │ ├── autonomous_agent.py
│ │ ├── planner_agent.py
│ │ ├── agent_interface.py
│ │ ├── team_orchestrator.py
│ │ ├── step_handler.py
│ │ ├── task_manager.py
│ │ ├── tools/
│ │ │ ├── calculator.py
│ │ │ ├── file_manager.py
│ │ │ ├── search_tool.py
│ │ │ └── tool_registry.py
│ │ ├── workflows/
│ │ │ ├── code_review_chain.py
│ │ │ ├── research_chain.py
│ │ │ ├── multi_agent_workflow.yaml
│ │ │ └── workflow_executor.py
│ │ └── security/ # policies; enforced by execution/executor.py
│ │ ├── tool_permissions.py # default-deny tool authorization (user perms ∩ agent profile)
│ │ ├── execution_policy.py # step, tool-call, token, cost and time budgets
│ │ └── approval_policy.py # human approval bound to argument hash
│ │
│ ├── cognition/
│ │ ├── cognitive_loop.py
│ │ ├── decision_policy.py
│ │ ├── planner.py
│ │ ├── reasoner.py
│ │ ├── state_interpreter.py
│ │ └── memory/
│ │ ├── long_term_memory.py
│ │ ├── short_term_memory.py
│ │ └── memory_manager.py
│ │
│ ├── execution/
│ │ ├── action_resolver.py
│ │ ├── controller.py
│ │ ├── error_handler.py
│ │ ├── executor.py
│ │ ├── threadpool.py # thread-offload helpers and limiters (no custom executor yet)
│ │ ├── job_scheduler.py # enqueue/schedule background jobs
│ │ └── background_worker.py # queue consumer / job dispatch loop
│ │
│ ├── llm/ # shared by RAG and agents
│ │ ├── client.py
│ │ ├── embeddings.py
│ │ ├── reranker.py
│ │ ├── model_loader.py
│ │ ├── cache.py
│ │ ├── dependencies.py # accessors for lifespan-created model clients
│ │ ├── config.py # LLM settings (LLM_ prefix, SecretStr keys)
│ │ ├── schemas.py # structured-output modes, capabilities, failure kinds
│ │ ├── structured_output.py # generate → validate → bounded retry (provider-independent)
│ │ └── prompts/
│ │ ├── system/
│ │ ├── tasks/
│ │ └── templates/
│ │
│ ├── guardrails/ # shared content-safety layer, used by RAG, agents and workers
│ │ ├── schemas.py # GuardrailPhase, GuardrailOutcome, GuardrailFinding, GuardrailVerdict
│ │ ├── service.py # run a phase's checks, decide outcome, enforce / log-only, fail closed
│ │ ├── policies.py # which checks apply to which surface and phase
│ │ ├── config.py # GUARDRAILS_ settings (mode, fail mode, thresholds, policy version)
│ │ ├── dependencies.py
│ │ ├── exceptions.py # GuardrailBlocked, GuardrailUnavailable
│ │ └── validators/ # deterministic in-house checks (regex, lists, length)
│ │
│ ├── providers/ # external-service boundaries (MCP, vector DB, queues, Lambda, guardrails, …)
│ │ ├── guardrails/client.py # managed guardrail API / validator-library adapters
│ │ ├── mcp/client.py
│ │ ├── vector_store/client.py
│ │ └── external/client.py
│ │
│ ├── evaluation/ # runtime-safe evaluation contracts only (no deepeval/ragas at import)
│ │ ├── schemas.py # EvaluationSample, RetrievedContext, AgentTrajectory, EvaluationResult
│ │ ├── registry.py # metric key → framework, class, version, required fields, result kind
│ │ └── adapters/
│ │ ├── deepeval_adapter.py # contracts ↔ LLMTestCase / ConversationalTestCase (lazy import)
│ │ └── ragas_adapter.py # contracts ↔ Ragas collections inputs / samples (lazy import)
│ │
│ └── common/
│ ├── schemas/
│ │ ├── base.py # CustomModel, UtcDatetime, shared annotated types
│ │ ├── responses.py # ErrorResponse and shared response shapes
│ │ └── pagination.py # Page[T], PageParams (models only; logic in src/pagination.py)
│ ├── dependencies.py
│ ├── logging.py
│ ├── retry.py
│ ├── timers.py
│ ├── serialization.py
│ ├── validation.py
│ └── security/
│ ├── headers.py # security-headers ASGI middleware (registered in src/middleware.py)
│ ├── rate_limiting.py # distributed limiter interface + key builders (backend 🟡)
│ ├── request_validation.py # body size/depth, content type, filenames, outbound URL (SSRF) policy
│ └── audit_schemas.py # AuditEvent, AuditActor, AuditTarget
│
├── tests/
│ ├── conftest.py # async client (httpx + ASGITransport), dependency-override fixtures
│ ├── auth/
│ ├── rag/{ingestion,retrieval,generation}/
│ ├── agents/
│ ├── cognition/
│ ├── execution/
│ ├── guardrails/ # fake checkers; deny / suppress / fail-closed / log-only paths
│ ├── evaluation/ # test_schemas.py, test_adapters.py, test_deterministic.py (no LLM calls)
│ ├── security/
│ │ ├── authentication/ # JWT/JWKS validation, throttling
│ │ ├── authorization/ # BOLA, BFLA, BOPLA, cross-tenant, service-to-service
│ │ ├── api/ # headers, CORS, limits, uploads, SSRF, injection, CSRF, rate limits
│ │ ├── rag/ # unauthorized/cross-tenant retrieval, ACL revocation, citation integrity
│ │ └── agents/ # tool authorization, argument BOLA, approval bypass, budgets, injection
│ ├── integration/
│ ├── e2e/
│ └── fixtures/
├── evaluation/ # offline quality evaluation (never imported by src/)
│ ├── README.md
│ ├── config/{metrics.yaml,thresholds.yaml}
│ ├── datasets/{rag,agents,conversations,golden,guardrails}/
│ ├── deepeval/{rag,agents,conversations,custom}/
│ ├── ragas/{rag,agents,custom}/
│ ├── deterministic/{citations,retrieval,structured_output,tool_calls}.py
│ ├── tracing/{collectors,normalization}.py
│ ├── runners/{offline,regression,report}.py
│ ├── fixtures/
│ └── reports/ # generated, git-ignored 🟡
├── data/{documents,knowledge_bases,agent_state,fixtures}/
├── scripts/
├── docs/
│ ├── architecture/
│ ├── api/
│ ├── development/
│ ├── decisions/
│ └── standards/
│ ├── ASYNC_EXECUTION.md # copied from this skill
│ ├── PYDANTIC_STANDARDS.md # copied from this skill
│ ├── API_CONVENTIONS.md # copied from this skill
│ ├── GUARDRAILS.md # copied from this skill
│ ├── AI_EVALUATION.md # copied from this skill
│ ├── API_SECURITY.md # copied from this skill
│ └── AGENT_SECURITY.md # copied from this skill
├── .env.example
├── .gitignore
├── pyproject.toml
├── README.md
└── CLAUDE.md
__init__.py files are omitted above for brevity. Rule 11 still applies.
3. Where does new code go?
| You are adding… | Put it in | Not in |
|---|---|---|
| A new business capability (orders, billing…) | src/<domain>/ using the example_domain/ layout | src/ root, common/ |
| An HTTP endpoint | The owning module's router.py (routes.py inside rag/), mounted via api/v1/router.py | main.py |
| Request/response models | <module>/schemas.py, inheriting CustomModel | src/models.py, direct BaseModel |
Shared base model, UtcDatetime, shared annotated types | common/schemas/base.py | Per-domain base classes |
| Error, message and page response shapes | common/schemas/{responses,pagination}.py | Each domain redefining them |
| Settings for a domain or component | <module>/config.py (own prefix, cached getter) | The global src/config.py |
| LLM-output schema for answer generation | rag/generation/schemas.py | rag/schemas.py (API) |
| LLM-output schema for one RAG stage (metadata, classification, expansion) | rag/<stage>/schemas.py, created when the stage is implemented | API schemas |
| Citation and answer checks against evidence | rag/generation/validation.py | Pydantic validators |
| Agent plan, tool-call union, final agent response | agents/structured_output.py | agents/schemas.py |
| A tool's argument model | That tool's module in agents/tools/ | A central args file |
| Structured-output generation, validation and retry flow | llm/structured_output.py + llm/schemas.py | Re-implemented in RAG or agents |
Evaluation contracts (EvaluationSample, EvaluationResult, …), metric registry | src/evaluation/{schemas,registry}.py | evaluation/, per-suite copies |
| Contract ↔ DeepEval / Ragas mapping | src/evaluation/adapters/ (lazy framework imports) | Suites building test cases ad hoc |
| Metric suites, G-Eval/DAG/custom metrics, judge wrappers | evaluation/{deepeval,ragas}/ | src/, tests/ |
| Deterministic metrics (retrieval P/R/MRR/NDCG, citations, schema, tool calls) | evaluation/deterministic/ | LLM judges |
| Evaluation datasets, metric/threshold config, runners, reports | evaluation/{datasets,config,runners,reports}/ | data/, tests/fixtures/ |
| Tests of the evaluation code | tests/evaluation/ | evaluation/ |
| DB models used by one domain | <domain>/models.py | src/models.py |
| Existence, ownership or uniqueness checks that need I/O | <domain>/dependencies.py (valid_<entity>_id) | Pydantic validators, repeated in each route |
current_principal, require_<permission>, scope dependencies | auth/dependencies.py | Each domain parsing tokens |
JWT/JWKS validation, claims → Principal | auth/tokens.py | auth/service.py, routes |
| RBAC/ABAC policy evaluation | auth/permissions.py | Routes, prompts |
| Password hashing (only if local accounts are approved 🟡) | auth/security.py | common/ |
Auth audit events; shared AuditEvent schema | auth/audit.py; common/security/audit_schemas.py | Free-form log lines |
| Security headers, rate limiting, request/upload/outbound-URL validation | common/security/{headers,rate_limiting,request_validation}.py | Per-domain copies |
| RAG access scope, document validation, post-retrieval policy | rag/security/ (scope resolved in rag/dependencies.py), applied as a store filter in retrieval | Prompts, the LLM, evidence_filter.py |
| Agent tool permissions, execution budgets, approval policy | agents/security/ (context built in agents/dependencies.py), enforced by execution/executor.py | Prompts, the planner, tools themselves |
| Security tests | tests/security/{authentication,authorization,api,rag,agents}/ | evaluation/ |
| Domain exceptions | <domain>/exceptions.py, subclassing src/exceptions.py | HTTPException in services |
| SQL queries | <domain>/service.py | Routers, dependencies |
| DB naming convention, engine, session dependency | src/database.py | Per-domain engines |
| Migrations 🟡 | alembic/ + alembic.ini at the root, only after approval | src/, the app lifespan |
| Lint/format script 🟡 | scripts/lint.sh | Ad-hoc commands in docs |
| Content-safety checks (input, context, tool, output) | Call src/guardrails/service.py at the checkpoints in GUARDRAILS.md §3 | Checks embedded in rag/ or agents/, system prompts |
| Deterministic safety validators (regex, lists) | guardrails/validators/ | common/validation.py |
| Managed guardrail or validator-library client | providers/guardrails/client.py | guardrails/, llm/ |
| Gateway Cedar guardrail policies 🟡 | Infrastructure-as-code, outside src/ (location needs approval) | src/ |
| Guardrail red-team / benign datasets | evaluation/datasets/guardrails/ | tests/ |
| Long-lived clients, pools and limiters | Created in main.py's lifespan and exposed through dependencies.py | Created per request |
| Document parsing and indexing logic | rag/extraction.py, rag/ingestion.py, rag/semantic/ | providers/, workers |
| Query routing / classification | rag/router/ | agents/ |
| Query rewriting, HyDE | rag/preprocessing/ | rag/retrieval/ |
| A retrieval or ranking stage | rag/retrieval/ | rag/generation/ |
| Answer context, citations, streaming | rag/generation/ | rag/chat/ |
| Conversation API / history | rag/chat/ | rag/ root |
| A new agent type | agents/<name>_agent.py implementing agent_interface.py | cognition/ |
| A tool an agent can call | agents/tools/ + register in tool_registry.py | common/ |
| A predefined multi-step workflow | agents/workflows/ | execution/ |
| Planning, reasoning, decision logic | cognition/ | agents/ |
| Agent memory | cognition/memory/ | rag/chat/storage.py |
| Running actions, control flow, execution errors | execution/{executor,controller,action_resolver,error_handler}.py | agents/ |
| Thread-offload helpers, per-dependency limiters | execution/threadpool.py | Ad-hoc in routes |
| Enqueueing or scheduling a background job | execution/job_scheduler.py | Routes calling the queue SDK directly |
| Queue consumer / job dispatch | execution/background_worker.py | A new app or rag/worker.py |
| Lambda handler (thin adapter) 🟡 | Same codebase. The location needs approval. It calls the existing service functions. | A separate repo or app with copied logic |
| Chat/completion, embedding or rerank model access | llm/ | rag/, agents/ |
| Prompt text | llm/prompts/{system,tasks,templates}/ | Scattered prompts.py files |
| MCP, vector DB, SQS, Lambda or third-party API client | providers/<kind>/client.py | rag/, agents/, llm/ |
| Logging, retry, timing, serialization | common/ | Domain packages |
| Unit tests | tests/<mirrored module>/ | Next to source |
| RAG/agent quality metrics, benchmarks | evaluation/ (layout in section 2) | tests/ |
| Sample docs, KB files, agent state | data/ | src/ |
| One-off CLIs (bulk upload, model training) | scripts/ | src/ |
| Engineering standards | docs/standards/ | README.md |
If none of these rows fit, ask the user before inventing a new top-level package.
4. RAG naming conventions 🟢
Use these names. The alternatives were rejected for the reasons given.
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 4k
- Forks
- 543
- Last commit
- Sep 2026
ahel review
K3info
injection (in AGENT_SECURITY.md)
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
production-fast-api-template- Source
- github.com/davepoon/buildwithclaude
github.com/davepoon/buildwithclaude
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonfastapi
Skill · fastapi
The pick for FastAPIintegration-fastapi
Skill · posthog
The pick for FastAPIobsidian-markdown
Skill · agricidaniel
The pick for Markdownmarkdown-formatter
Skill · nvidia
The pick for Markdown