Production FastAPI Template: Rules

SkillCloud & infra

Structure and execution rules for a unified, production-ready backend combining domain-oriented FastAPI, Retrieval-Augmented Generation (RAG) and Agentic AI in one codebase on AWS ECS Fargate. Use when scaffolding a new backend; adding a module, API domain, RAG component, agent, tool, workflow, LLM/provider integration or test; deciding where a file belongs; writing async/sync routes, services or clients; offloading blocking or CPU-heavy work (thread pools, process pools, AWS Lambda); designing background jobs, queues (SQS, Celery), workers and scaling; or writing Pydantic schemas, settings (pydantic-settings), serialization, structured LLM outputs, citation validation, agent plans or tool-call arguments; or writing FastAPI dependencies (request validation, auth, access scopes, chaining), REST paths, response models, OpenAPI docs, error responses, database naming and SQL queries, migrations, API tests (async client, dependency overrides) or lint/format setup (ruff); or adding content-safety guardrails (input/output checks, PII, prompt injection, toxicity, Bedrock Guardrails, AgentCore Policy/Gateway Cedar policies, Guardrails AI validators), fail-closed handling and guardrail evaluation; or evaluating RAG, retrieval stages, structured outputs, citations, agents, multi-agent workflows or conversations (DeepEval, Ragas, LLM-as-judge, G-Eval, golden datasets, regression tiers, thresholds, judge calibration); or securing the API (authentication, JWT/JWKS, OAuth 2.0/OIDC, RBAC/ABAC, BOLA/BFLA, tenant isolation, input/upload/SSRF validation, security headers, CORS, CSRF, rate limiting, audit logging, ECS/IAM/secrets hardening, CI security scanning, security tests) or LLM/agent security (OWASP LLM Top 10, OWASP Agentic Top 10, prompt injection, tool authorization, human approval, excessive agency).

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Production FastAPI Template: Rules skill

What this skill tells your AI

The instructions your AI receives, as published by davepoon/buildwithclaude in plugins/production-fast-api-template/skills/production-fast-api-template/SKILL.md and read by ahel’s review.

These rules define where code lives and how it executes in a unified FastAPI + RAG + Agentic AI backend. Apply them whenever you generate or modify code in a project built from this template.

Companion references. Read the relevant file before writing code:

FileRead before writing…Summarized in
ASYNC_EXECUTION.mdAny route, service, client, agent tool, worker or background jobSection 5
PYDANTIC_STANDARDS.mdAny schema, settings class, serialization, LLM call with structured output, citation check, agent plan or tool argument modelSection 6
API_CONVENTIONS.mdAny dependency, route signature, auth/permission check, OpenAPI metadata, DB model or query, migration, API test or lint setupSection 7
GUARDRAILS.mdAny user-facing LLM call, RAG answer, ingestion of untrusted documents, agent tool execution, gateway target or content-safety checkSection 8
AI_EVALUATION.mdAny evaluation contract, metric, dataset, judge, runner, report or threshold; any new RAG stage, agent, tool or conversation feature that needs quality coverageSection 9
API_SECURITY.mdAny authentication, token, permission, tenant or ownership check; input/upload/outbound-URL handling; response headers, CORS, errors; rate limits; audit events; ECS/IAM/secrets; CI security; security testsSection 10
AGENT_SECURITY.mdAny agent, tool, tool registry entry, approval flow, memory write, MCP connector or LLM-output sinkSection 10

Status labels:

  • 🟢 Convention. A proposed rule. Follow it by default.
  • 🟡 Needs approval. An open team decision. Surface the options to the user and never pick one silently.

These rules don't choose agent frameworks, LLM providers, models, vector stores or a queue implementation. If a rule doesn't fit a real need, raise it with the user instead of silently working around it.


1. Core structure rules 🟢

  1. One application. FastAPI, RAG and Agentic AI are not separate apps. There is exactly one src/ package, one FastAPI entry point (src/main.py) and one shared infrastructure layer.
  2. Never create a second entry point. Don't add a second src/, a separate app folder (rag_app/, agents_service/, worker_app/) or an extra FastAPI() instance.
  3. Background workers and Lambda functions are extra entry points into the same codebase. They are not separate applications, and they call the same service functions the API calls.
  4. One central API router. All routers mount through src/api/v1/router.py. Domain, RAG and agent routers never mount themselves on the app.
  5. Domain-oriented modules. Business code lives inside its domain package. src/'s root holds only application-wide concerns.
  6. No duplicated shared layers. Model clients go in src/llm/, content-safety checks in src/guardrails/ and external-service clients in src/providers/, each exactly once. RAG and agents consume them and never define their own copies (no rag/llm_client.py, no agents/vector_db.py).
  7. Provider-neutral naming. Name files by responsibility, not vendor (llm/embeddings.py, not openai_embeddings.py). Vendor specifics stay behind src/providers/ or src/llm/.
  8. common/ is for genuinely reusable, domain-agnostic code only. Never put business logic there.
  9. Keep agents, reasoning and execution apart. Agent definitions go in agents/, thinking in cognition/ and doing, scheduling and offloading in execution/.
  10. Tests mirror src/. Quality evaluation (datasets, DeepEval/Ragas suites, deterministic metrics, runners, reports) lives in the top-level evaluation/ directory, separate from tests/. src/evaluation/ holds only framework-neutral contracts, the metric registry and adapters, with lazy framework imports (AI_EVALUATION.md §1).
  11. Every Python directory under src/ has an __init__.py. Non-Python directories that must survive in Git (prompt folders, data, empty test folders) get a .gitkeep.
  12. Explicit module imports across packages, e.g. from src.auth import constants as auth_constants.
  13. Reuse existing files before adding new ones. Extend execution/background_worker.py rather than creating execution/worker2.py or rag/worker.py.

2. Canonical structure

<project-name>/
├── src/
│   ├── __init__.py
│   ├── main.py              # the single FastAPI entry point (lifespan creates shared clients/limiters)
│   ├── config.py            # application-wide configuration
│   ├── constants.py
│   ├── exceptions.py        # base domain exception + handlers rendering ErrorResponse
│   ├── middleware.py
│   ├── database.py          # shared metadata (naming convention), engine/session factory, get_db_session
│   ├── models.py            # genuinely shared models only
│   ├── pagination.py
│   │
│   ├── api/
│   │   └── v1/
│   │       └── router.py    # aggregates domain, RAG and agent routers
│   │
│   ├── auth/                # reference domain module
│   │   ├── router.py
│   │   ├── schemas.py       # the domain's own Pydantic API/internal schemas
│   │   ├── models.py        # persistence (ORM) models, not Pydantic
│   │   ├── dependencies.py  # current_principal, require_<permission>, scope dependencies
│   │   ├── config.py        # the domain's own pydantic-settings class (AUTH_ prefix: issuers, audiences, JWKS)
│   │   ├── constants.py     # permission and scope names
│   │   ├── exceptions.py
│   │   ├── service.py
│   │   ├── tokens.py        # JWT/JWKS validation, token types, claims → Principal
│   │   ├── permissions.py   # RBAC/ABAC policy evaluation (default deny)
│   │   ├── security.py      # password hashing, only if local accounts are approved 🟡
│   │   ├── audit.py         # auth audit events (login, token rejected, session revoked)
│   │   └── utils.py
│   │
│   ├── example_domain/      # copy this layout for every new business domain
│   │   ├── router.py
│   │   ├── schemas.py
│   │   ├── models.py
│   │   ├── dependencies.py  # valid_<entity>_id and other I/O-backed validation
│   │   ├── constants.py
│   │   ├── exceptions.py
│   │   ├── service.py
│   │   └── utils.py
│   │
│   ├── rag/
│   │   ├── routes.py        # "routes", not "router": avoids clashing with rag/router/
│   │   ├── schemas.py       # RAG API schemas + cross-stage contracts (RetrievedEvidence)
│   │   ├── service.py       # container/lifecycle of long-lived RAG services
│   │   ├── dependencies.py
│   │   ├── config.py        # RAG settings (RAG_ prefix)
│   │   ├── constants.py
│   │   ├── exceptions.py
│   │   ├── extraction.py    # structure-preserving document extraction
│   │   ├── ingestion.py     # coordinates processing and indexing (run by background jobs)
│   │   ├── semantic/
│   │   │   ├── metadata.py
│   │   │   ├── segmentation.py
│   │   │   └── chunks.py
│   │   ├── router/
│   │   │   ├── classifier.py
│   │   │   ├── fusion.py
│   │   │   └── evaluation.py
│   │   ├── preprocessing/
│   │   │   ├── query_expansion.py
│   │   │   └── hyde.py
│   │   ├── retrieval/
│   │   │   ├── semantic_search.py
│   │   │   ├── bm25.py
│   │   │   ├── rrf.py
│   │   │   ├── mmr.py
│   │   │   ├── evidence_filter.py
│   │   │   └── hybrid_retriever.py
│   │   ├── generation/
│   │   │   ├── context_builder.py
│   │   │   ├── citations.py
│   │   │   ├── streaming.py
│   │   │   ├── schemas.py        # LLM-output contracts: GeneratedAnswer, Citation
│   │   │   ├── validation.py     # application-level answer/citation validation
│   │   │   └── service.py
│   │   ├── chat/
│   │   │   ├── routes.py
│   │   │   ├── schemas.py
│   │   │   ├── service.py
│   │   │   └── storage.py
│   │   └── security/
│   │       ├── access_control.py     # AccessScope model + vector-store/DB filter translation
│   │       ├── document_validation.py  # ingestion file checks (type, signature, size, archives)
│   │       └── retrieval_policy.py   # post-retrieval ACL re-check, trust labels, provenance
│   │
│   ├── agents/
│   │   ├── router.py
│   │   ├── schemas.py            # AgentTask, ToolResult, AgentExecutionState
│   │   ├── structured_output.py  # LLM-produced: ExecutionPlan, ToolCall union, FinalAgentResponse
│   │   ├── dependencies.py
│   │   ├── config.py             # agent settings (AGENTS_ prefix)
│   │   ├── constants.py
│   │   ├── exceptions.py
│   │   ├── service.py
│   │   ├── base_agent.py
│   │   ├── autonomous_agent.py
│   │   ├── planner_agent.py
│   │   ├── agent_interface.py
│   │   ├── team_orchestrator.py
│   │   ├── step_handler.py
│   │   ├── task_manager.py
│   │   ├── tools/
│   │   │   ├── calculator.py
│   │   │   ├── file_manager.py
│   │   │   ├── search_tool.py
│   │   │   └── tool_registry.py
│   │   ├── workflows/
│   │   │   ├── code_review_chain.py
│   │   │   ├── research_chain.py
│   │   │   ├── multi_agent_workflow.yaml
│   │   │   └── workflow_executor.py
│   │   └── security/                 # policies; enforced by execution/executor.py
│   │       ├── tool_permissions.py   # default-deny tool authorization (user perms ∩ agent profile)
│   │       ├── execution_policy.py   # step, tool-call, token, cost and time budgets
│   │       └── approval_policy.py    # human approval bound to argument hash
│   │
│   ├── cognition/
│   │   ├── cognitive_loop.py
│   │   ├── decision_policy.py
│   │   ├── planner.py
│   │   ├── reasoner.py
│   │   ├── state_interpreter.py
│   │   └── memory/
│   │       ├── long_term_memory.py
│   │       ├── short_term_memory.py
│   │       └── memory_manager.py
│   │
│   ├── execution/
│   │   ├── action_resolver.py
│   │   ├── controller.py
│   │   ├── error_handler.py
│   │   ├── executor.py
│   │   ├── threadpool.py         # thread-offload helpers and limiters (no custom executor yet)
│   │   ├── job_scheduler.py      # enqueue/schedule background jobs
│   │   └── background_worker.py  # queue consumer / job dispatch loop
│   │
│   ├── llm/                 # shared by RAG and agents
│   │   ├── client.py
│   │   ├── embeddings.py
│   │   ├── reranker.py
│   │   ├── model_loader.py
│   │   ├── cache.py
│   │   ├── dependencies.py       # accessors for lifespan-created model clients
│   │   ├── config.py             # LLM settings (LLM_ prefix, SecretStr keys)
│   │   ├── schemas.py            # structured-output modes, capabilities, failure kinds
│   │   ├── structured_output.py  # generate → validate → bounded retry (provider-independent)
│   │   └── prompts/
│   │       ├── system/
│   │       ├── tasks/
│   │       └── templates/
│   │
│   ├── guardrails/          # shared content-safety layer, used by RAG, agents and workers
│   │   ├── schemas.py            # GuardrailPhase, GuardrailOutcome, GuardrailFinding, GuardrailVerdict
│   │   ├── service.py            # run a phase's checks, decide outcome, enforce / log-only, fail closed
│   │   ├── policies.py           # which checks apply to which surface and phase
│   │   ├── config.py             # GUARDRAILS_ settings (mode, fail mode, thresholds, policy version)
│   │   ├── dependencies.py
│   │   ├── exceptions.py         # GuardrailBlocked, GuardrailUnavailable
│   │   └── validators/           # deterministic in-house checks (regex, lists, length)
│   │
│   ├── providers/           # external-service boundaries (MCP, vector DB, queues, Lambda, guardrails, …)
│   │   ├── guardrails/client.py  # managed guardrail API / validator-library adapters
│   │   ├── mcp/client.py
│   │   ├── vector_store/client.py
│   │   └── external/client.py
│   │
│   ├── evaluation/          # runtime-safe evaluation contracts only (no deepeval/ragas at import)
│   │   ├── schemas.py            # EvaluationSample, RetrievedContext, AgentTrajectory, EvaluationResult
│   │   ├── registry.py           # metric key → framework, class, version, required fields, result kind
│   │   └── adapters/
│   │       ├── deepeval_adapter.py   # contracts ↔ LLMTestCase / ConversationalTestCase (lazy import)
│   │       └── ragas_adapter.py      # contracts ↔ Ragas collections inputs / samples (lazy import)
│   │
│   └── common/
│       ├── schemas/
│       │   ├── base.py           # CustomModel, UtcDatetime, shared annotated types
│       │   ├── responses.py      # ErrorResponse and shared response shapes
│       │   └── pagination.py     # Page[T], PageParams (models only; logic in src/pagination.py)
│       ├── dependencies.py
│       ├── logging.py
│       ├── retry.py
│       ├── timers.py
│       ├── serialization.py
│       ├── validation.py
│       └── security/
│           ├── headers.py            # security-headers ASGI middleware (registered in src/middleware.py)
│           ├── rate_limiting.py      # distributed limiter interface + key builders (backend 🟡)
│           ├── request_validation.py # body size/depth, content type, filenames, outbound URL (SSRF) policy
│           └── audit_schemas.py      # AuditEvent, AuditActor, AuditTarget
│
├── tests/
│   ├── conftest.py          # async client (httpx + ASGITransport), dependency-override fixtures
│   ├── auth/
│   ├── rag/{ingestion,retrieval,generation}/
│   ├── agents/
│   ├── cognition/
│   ├── execution/
│   ├── guardrails/          # fake checkers; deny / suppress / fail-closed / log-only paths
│   ├── evaluation/          # test_schemas.py, test_adapters.py, test_deterministic.py (no LLM calls)
│   ├── security/
│   │   ├── authentication/  # JWT/JWKS validation, throttling
│   │   ├── authorization/   # BOLA, BFLA, BOPLA, cross-tenant, service-to-service
│   │   ├── api/             # headers, CORS, limits, uploads, SSRF, injection, CSRF, rate limits
│   │   ├── rag/             # unauthorized/cross-tenant retrieval, ACL revocation, citation integrity
│   │   └── agents/          # tool authorization, argument BOLA, approval bypass, budgets, injection
│   ├── integration/
│   ├── e2e/
│   └── fixtures/
├── evaluation/              # offline quality evaluation (never imported by src/)
│   ├── README.md
│   ├── config/{metrics.yaml,thresholds.yaml}
│   ├── datasets/{rag,agents,conversations,golden,guardrails}/
│   ├── deepeval/{rag,agents,conversations,custom}/
│   ├── ragas/{rag,agents,custom}/
│   ├── deterministic/{citations,retrieval,structured_output,tool_calls}.py
│   ├── tracing/{collectors,normalization}.py
│   ├── runners/{offline,regression,report}.py
│   ├── fixtures/
│   └── reports/             # generated, git-ignored 🟡
├── data/{documents,knowledge_bases,agent_state,fixtures}/
├── scripts/
├── docs/
│   ├── architecture/
│   ├── api/
│   ├── development/
│   ├── decisions/
│   └── standards/
│       ├── ASYNC_EXECUTION.md      # copied from this skill
│       ├── PYDANTIC_STANDARDS.md   # copied from this skill
│       ├── API_CONVENTIONS.md      # copied from this skill
│       ├── GUARDRAILS.md           # copied from this skill
│       ├── AI_EVALUATION.md        # copied from this skill
│       ├── API_SECURITY.md         # copied from this skill
│       └── AGENT_SECURITY.md       # copied from this skill
├── .env.example
├── .gitignore
├── pyproject.toml
├── README.md
└── CLAUDE.md

__init__.py files are omitted above for brevity. Rule 11 still applies.


3. Where does new code go?

You are adding…Put it inNot in
A new business capability (orders, billing…)src/<domain>/ using the example_domain/ layoutsrc/ root, common/
An HTTP endpointThe owning module's router.py (routes.py inside rag/), mounted via api/v1/router.pymain.py
Request/response models<module>/schemas.py, inheriting CustomModelsrc/models.py, direct BaseModel
Shared base model, UtcDatetime, shared annotated typescommon/schemas/base.pyPer-domain base classes
Error, message and page response shapescommon/schemas/{responses,pagination}.pyEach domain redefining them
Settings for a domain or component<module>/config.py (own prefix, cached getter)The global src/config.py
LLM-output schema for answer generationrag/generation/schemas.pyrag/schemas.py (API)
LLM-output schema for one RAG stage (metadata, classification, expansion)rag/<stage>/schemas.py, created when the stage is implementedAPI schemas
Citation and answer checks against evidencerag/generation/validation.pyPydantic validators
Agent plan, tool-call union, final agent responseagents/structured_output.pyagents/schemas.py
A tool's argument modelThat tool's module in agents/tools/A central args file
Structured-output generation, validation and retry flowllm/structured_output.py + llm/schemas.pyRe-implemented in RAG or agents
Evaluation contracts (EvaluationSample, EvaluationResult, …), metric registrysrc/evaluation/{schemas,registry}.pyevaluation/, per-suite copies
Contract ↔ DeepEval / Ragas mappingsrc/evaluation/adapters/ (lazy framework imports)Suites building test cases ad hoc
Metric suites, G-Eval/DAG/custom metrics, judge wrappersevaluation/{deepeval,ragas}/src/, tests/
Deterministic metrics (retrieval P/R/MRR/NDCG, citations, schema, tool calls)evaluation/deterministic/LLM judges
Evaluation datasets, metric/threshold config, runners, reportsevaluation/{datasets,config,runners,reports}/data/, tests/fixtures/
Tests of the evaluation codetests/evaluation/evaluation/
DB models used by one domain<domain>/models.pysrc/models.py
Existence, ownership or uniqueness checks that need I/O<domain>/dependencies.py (valid_<entity>_id)Pydantic validators, repeated in each route
current_principal, require_<permission>, scope dependenciesauth/dependencies.pyEach domain parsing tokens
JWT/JWKS validation, claims → Principalauth/tokens.pyauth/service.py, routes
RBAC/ABAC policy evaluationauth/permissions.pyRoutes, prompts
Password hashing (only if local accounts are approved 🟡)auth/security.pycommon/
Auth audit events; shared AuditEvent schemaauth/audit.py; common/security/audit_schemas.pyFree-form log lines
Security headers, rate limiting, request/upload/outbound-URL validationcommon/security/{headers,rate_limiting,request_validation}.pyPer-domain copies
RAG access scope, document validation, post-retrieval policyrag/security/ (scope resolved in rag/dependencies.py), applied as a store filter in retrievalPrompts, the LLM, evidence_filter.py
Agent tool permissions, execution budgets, approval policyagents/security/ (context built in agents/dependencies.py), enforced by execution/executor.pyPrompts, the planner, tools themselves
Security teststests/security/{authentication,authorization,api,rag,agents}/evaluation/
Domain exceptions<domain>/exceptions.py, subclassing src/exceptions.pyHTTPException in services
SQL queries<domain>/service.pyRouters, dependencies
DB naming convention, engine, session dependencysrc/database.pyPer-domain engines
Migrations 🟡alembic/ + alembic.ini at the root, only after approvalsrc/, the app lifespan
Lint/format script 🟡scripts/lint.shAd-hoc commands in docs
Content-safety checks (input, context, tool, output)Call src/guardrails/service.py at the checkpoints in GUARDRAILS.md §3Checks embedded in rag/ or agents/, system prompts
Deterministic safety validators (regex, lists)guardrails/validators/common/validation.py
Managed guardrail or validator-library clientproviders/guardrails/client.pyguardrails/, llm/
Gateway Cedar guardrail policies 🟡Infrastructure-as-code, outside src/ (location needs approval)src/
Guardrail red-team / benign datasetsevaluation/datasets/guardrails/tests/
Long-lived clients, pools and limitersCreated in main.py's lifespan and exposed through dependencies.pyCreated per request
Document parsing and indexing logicrag/extraction.py, rag/ingestion.py, rag/semantic/providers/, workers
Query routing / classificationrag/router/agents/
Query rewriting, HyDErag/preprocessing/rag/retrieval/
A retrieval or ranking stagerag/retrieval/rag/generation/
Answer context, citations, streamingrag/generation/rag/chat/
Conversation API / historyrag/chat/rag/ root
A new agent typeagents/<name>_agent.py implementing agent_interface.pycognition/
A tool an agent can callagents/tools/ + register in tool_registry.pycommon/
A predefined multi-step workflowagents/workflows/execution/
Planning, reasoning, decision logiccognition/agents/
Agent memorycognition/memory/rag/chat/storage.py
Running actions, control flow, execution errorsexecution/{executor,controller,action_resolver,error_handler}.pyagents/
Thread-offload helpers, per-dependency limitersexecution/threadpool.pyAd-hoc in routes
Enqueueing or scheduling a background jobexecution/job_scheduler.pyRoutes calling the queue SDK directly
Queue consumer / job dispatchexecution/background_worker.pyA new app or rag/worker.py
Lambda handler (thin adapter) 🟡Same codebase. The location needs approval. It calls the existing service functions.A separate repo or app with copied logic
Chat/completion, embedding or rerank model accessllm/rag/, agents/
Prompt textllm/prompts/{system,tasks,templates}/Scattered prompts.py files
MCP, vector DB, SQS, Lambda or third-party API clientproviders/<kind>/client.pyrag/, agents/, llm/
Logging, retry, timing, serializationcommon/Domain packages
Unit teststests/<mirrored module>/Next to source
RAG/agent quality metrics, benchmarksevaluation/ (layout in section 2)tests/
Sample docs, KB files, agent statedata/src/
One-off CLIs (bulk upload, model training)scripts/src/
Engineering standardsdocs/standards/README.md

If none of these rows fit, ask the user before inventing a new top-level package.


4. RAG naming conventions 🟢

Use these names. The alternatives were rejected for the reasons given.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
4k
Forks
543
Last commit
Sep 2026

ahel review

  • K3info
    injection (in AGENT_SECURITY.md)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
production-fast-api-template
Source
github.com/davepoon/buildwithclaude