Writing Streamlit apps for the PostHog sandbox
SkillMediaWriting-streamlit-apps is a skill that teaches an AI agent how to write Streamlit app source code that runs well in a PostHog sandbox. It covers the posthog_apps.query() bridge for reading PostHog data, the packages available in the sandbox image, caching and session state across Streamlit reruns, layout and chart patterns, and th
Available today. Use it from your connected AI after setup.
No other account needed.
Have an agent that can load skills, and a PostHog project with Streamlit apps you want to build or debug.
Then ask your AI: use the Writing Streamlit apps for the PostHog sandbox skill
What your AI can do with it
- Write app.py source that runs in the sandboxed Streamlit 1.31 runtime
- Read PostHog data via the posthog_apps.query() bridge, which returns a pandas DataFrame
- Cache bridge calls with st.cache_data so widget interactions don't re-fire queries
- Use st.session_state for values that must survive Streamlit reruns
- Avoid interpolating free-text widget values into HogQL to prevent access leaks
- Handle RuntimeError from failed queries and render errors with st.error
Getting started
- Have an agent that can load skills, and a PostHog project with Streamlit apps you want to build or debug.
- Add the writing-streamlit-apps skill to the agent's available skills.
- Ask the agent to write or debug a Streamlit app that shows PostHog data, or invoke it when a query inside an app fails.
- Test queries in the SQL editor where real errors are visible, since the sandbox bridge returns generic error messages.
What this skill tells your AI
The instructions your AI receives, as published by posthog/posthog in products/streamlit_apps/skills/writing-streamlit-apps/SKILL.md and read by ahel’s review.
The source you write becomes app.py at the root of a sandboxed Streamlit 1.31 runtime.
Deployment mechanics (create/start/share) are the managing-streamlit-apps skill; this one is about the code.
Reading PostHog data: posthog_apps.query()
The one and only data door is the in-sandbox bridge:
import posthog_apps
df = posthog_apps.query("SELECT event, count() FROM events GROUP BY event LIMIT 10")
- Takes a HogQL string, returns a pandas DataFrame.
- Raises
RuntimeErroron failure. The message is deliberately generic ("Query execution failed") — the bridge does not return query internals to the sandbox, so you cannot diagnose a bad query from inside the app. Catch it and render withst.error(str(e))so viewers get a message instead of a stack trace, and test queries in the SQL editor where real errors are visible. import posthogdoes NOT exist in the sandbox — the module isposthog_apps, deliberately distinct from the posthog-python SDK's name.- The bridge is pre-authenticated to the app's project; user code never sees a token, and there is nothing to configure.
- Queries run with server-side caps (30 s execution, 256 MB memory) that don't scale with sandbox sizing — so bound time ranges,
LIMITresults, and aggregate in HogQL rather than pulling raw events into pandas.
Design for Streamlit's rerun model
Streamlit reruns the whole script top to bottom on every widget interaction. Two consequences:
-
Cache every bridge call — uncached, one slider drag re-fires every query:
@st.cache_data(ttl=300, show_spinner="Running query...") def run_query(hogql: str) -> pd.DataFrame: return posthog_apps.query(hogql) -
Use
st.session_statefor anything that must survive reruns — accumulated selections, pagination cursors, "last refreshed" stamps. Module-level variables reset on every interaction.
Widgets drive parameters naturally, but never interpolate a free-text widget value into HogQL.
The bridge runs your query with the version author's data access, and anyone who can view the app drives those widgets — a raw st.text_input spliced into a query hands viewers the author's access to write their own.
Constrain the input instead: pick from a fixed list you control (st.selectbox over known values), or coerce to a type that can't carry SQL (int(days), a date from st.date_input), and validate before it reaches the query.
Layout and charts
st.set_page_config(page_title=..., layout="wide")first — the default narrow layout wastes most of the screen for data apps.- Structure with
st.columnsfor side-by-side metrics,st.tabsfor alternate views,st.expanderfor detail sections;st.metricfor headline numbers. - Charts:
st.plotly_chart(fig, use_container_width=True)with plotly express is the reliable default;st.dataframe(df, use_container_width=True)for tables. matplotlib/seaborn also work viast.pyplot.
What's installed
The image ships Python 3.11 with: streamlit 1.31, pandas, numpy, polars, plotly, matplotlib, seaborn, scipy, scikit-learn, pyarrow, duckdb, requests, beautifulsoup4, lxml, sqlalchemy, aiohttp.
There is no way to add dependencies: the sandbox never runs pip (a deliberate security posture — no arbitrary package code at boot), and a requirements.txt in an uploaded zip is tolerated but dropped. Only import what's listed above.
Structure and runtime constraints
app.pyis the entry point. Via the MCP set-source flow your source ISapp.py. Helper modules and data files can ship next to it through thefiles(text) andassets(base64) maps;import utilsworks because Streamlit puts the app directory onsys.path, but the process does not run inside that directory, so open data files viaPath(__file__).parent / "data/events.parquet", never a bare relative path. Keep bundled data small: it is sent inline as JSON on every set-source call.- The sandbox is ephemeral: anything written to disk disappears on stop/restart. Don't build state on files; recompute from queries (with caching) or hold it in
st.session_state. - Your code runs as an unprivileged user; there's no posthog SDK, no way to pass your own environment variables or secrets to the app, and no expectation of general network egress. Don't read from
os.environ— anything there belongs to the sandbox runtime, not your app. Design aroundposthog_apps.query()as the data source.
A minimal well-shaped app
import pandas as pd
import plotly.express as px
import posthog_apps
import streamlit as st
st.set_page_config(page_title="Events overview", layout="wide")
st.title("Events overview")
@st.cache_data(ttl=300, show_spinner="Running query...")
def run_query(hogql: str) -> pd.DataFrame:
return posthog_apps.query(hogql)
# A slider is bounded and coerced to int, so it is safe to interpolate.
days = int(st.slider("Days to show", 1, 30, 7))
try:
daily = run_query(
f"""
SELECT toDate(timestamp) AS day, count() AS events
FROM events
WHERE timestamp >= now() - INTERVAL {days} DAY
GROUP BY day ORDER BY day
"""
)
st.plotly_chart(px.bar(daily, x="day", y="events"), use_container_width=True)
except RuntimeError as e:
st.error(str(e))
Signals
- GitHub stars
- 40k
- Forks
- 3k
- Last commit
- Sep 2026
Others that do the same job
Questions
- How does an app read PostHog data?
- Through the posthog_apps.query() bridge. It takes a HogQL string and returns a pandas DataFrame. The bridge is pre-authenticated to the app's project, so no token or configuration is needed. Note that import posthog does not exist in the sandbox.
- Why did my query fail with a generic error?
- The bridge raises RuntimeError with a deliberately generic message and does not return query internals to the sandbox. Test the query in the SQL editor where real errors are visible, and catch the error in the app to render it with st.error.
Advanced
- Item type
- skill
- Key
writing-streamlit-apps- Source
- github.com/posthog/posthog