Server Configuration
ThinkingSDK server settings control how exceptions are processed, how the autofix pipeline operates, and how billing works.
Environment Variables
The server reads configuration from environment variables. Set these in your .env file or deployment environment.
Core Settings
# Database
DATABASE_URL=postgresql://user:password@host:5432/thinkingsdk
# Server
APP_URL=https://your-domain.com
PORT=8000
# Authentication
OAUTH_CLIENT_ID=your_google_oauth_client_id
OAUTH_CLIENT_SECRET=your_google_oauth_secret
SESSION_SECRET=your_random_session_secret
GitHub Integration
# GitHub OAuth App credentials (for connecting user repos)
GITHUB_CLIENT_ID=your_github_oauth_client_id
GITHUB_CLIENT_SECRET=your_github_oauth_secret
When a user connects their GitHub account, ThinkingSDK can:
- Create issues for detected exceptions
- Clone repositories to generate fixes
- Open pull requests with AI-generated patches
- Update issues with fix progress
AI / LLM Configuration
# Anthropic (Claude) — primary fix generation engine
ANTHROPIC_API_KEY=sk-ant-your-key
# OpenAI (fallback)
OPENAI_API_KEY=sk-your-key
ThinkingSDK uses Claude as the primary engine for:
- Code analysis and root cause identification
- Fix generation with full repository context
- Fix quality evaluation (judge)
- Test generation
Stripe Billing
STRIPE_SECRET_KEY=sk_test_your_key
STRIPE_PUBLISHABLE_KEY=pk_test_your_key
STRIPE_WEBHOOK_SECRET=whsec_your_secret
Autofix Pipeline
The autofix pipeline turns captured exceptions into fixes. A worker claims one exception group at a time and runs it through a short list of steps.
Pipeline Phases
| Phase | What It Does |
|---|---|
| Phase 1: Intake | Groups exceptions, checks org tier and usage limits, queues new groups |
| Phase 2: Setup | Creates GitHub issue, clones repository, sets up workspace |
| Phase 3: Fix Generation | Analyzes code, generates fix, evaluates with judge, runs tests |
| Phase 4: PR Creation | Creates pull request, updates issue with PR link |
Pipeline Behavior by Tier
| Tier | Crash Capture | Time Capsule | Dashboard | Autofix + PR |
|---|---|---|---|---|
| Free | Yes | Yes | Yes | No |
| Contact Founders | Yes | Yes | Yes | Yes |
Free tier users get full crash capture, Time Capsule viewing, and dashboard access. The autofix pipeline (AI fix generation + PR creation) requires a paid subscription.
States and Stages
Each exception group has a state (what the worker does with it) and a stage (how far it got):
queued -> running -> succeeded
-> queued retryable error, retried with exponential backoff
-> waiting needs you, e.g. connect GitHub, then press Retry Fix
-> failed permanent error or retries used up
Stages, in order:
with a repository: issue -> workspace -> fix -> pull_request
without a repository: analysis -> suggestion
A retry resumes after the last stage whose result is saved: an existing GitHub issue is never created twice and a group that already has a pull request is not fixed again.
Retries
- Every run of a group is one attempt. A group is retried automatically up to 6 attempts, waiting 1, 2, 4 ... minutes between them (at most 1 hour; GitHub rate limits wait for the limit to reset).
- Inside one attempt the fix stage tries up to 3 times, feeding the judge's feedback or test failures into the next try.
- A failed group is retried automatically (at most twice, at most once a day) when the exception keeps happening and its occurrence count doubles.
- Retry Fix on the exception page moves a waiting or failed group back to the queue with a fresh attempt budget.
Remote Client Configuration
Control SDK behavior server-side without redeploying client applications.
API Endpoints
GET /api/v1/client-config — Fetch current config
PUT /api/v1/client-config — Update config
GET /api/v1/client-config/history — View change history
DELETE /api/v1/client-config — Reset to defaults
Example: Change Sample Rate Remotely
curl -X PUT https://api.thinkingsdk.ai/api/v1/client-config \
-H "X-THINKINGSDK-KEY: tsdk_your_key" \
-H "Content-Type: application/json" \
-d '{
"instrumentation": {
"sample_rate": 0.1
}
}'
All connected clients with remote_config_enabled: true will pick up this change within their poll interval (default: 5 minutes).
Configurable Remotely
- Sample rate
- Ignore patterns (files/functions to skip)
- Batch size and timing
- Feature toggles (capture_memory, capture_source_lines)
- Custom sampling rules
API Reference
Event Ingestion
POST /ingest
Accepts exception events from the SDK client. Supports batch and single events, gzip compression.
Headers (as sent by the Python client 0.1.6):
- X-Session-Token: sess_...: session from POST /auth/session; the client falls back to X-THINKINGSDK-KEY: tsdk_your_api_key when it cannot get one
- Content-Type: application/json
- Content-Encoding: gzip: bodies over 1 KiB
- User-Agent: ThinkingSDK-Python/0.1.6 (Python 3.12.11; Darwin)
- X-SDK-Language: Python
- X-SDK-Version: 0.1.6
Repeats of a crash within the client's 15 minute deduplication window arrive as one deduplicated_pattern event; its data.frequency is added to the group's occurrence count.
Response:
{
"status": "ok",
"events_received": 5,
"events_stored": 5
}
Exception Flow
GET /api/event-flow/{event_id}
Returns the complete processing flow for an exception: event details, exception group, state transitions, fix attempts, and fix suggestions.
Crash Context (Time Capsule)
GET /api/crash-context/{event_id}
Returns the full frozen crash state: stack frames with source code, local/global variables at every frame, breadcrumbs, system snapshot, and AI fix data.
Billing Usage
GET /api/billing/usage
Returns current event count, fix attempt count, and tier limits.
Manual Retry
POST /api/manual-retry/{exception_group_id}
Manually trigger a retry for a failed fix attempt. Resets the iteration counter and starts a fresh fix attempt.
Deployment
ThinkingSDK server runs on any platform that supports Python and PostgreSQL.
Railway (Recommended)
# Procfile
web: uvicorn webapp.server:app --host 0.0.0.0 --port $PORT
Docker
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
CMD ["uvicorn", "webapp.server:app", "--host", "0.0.0.0", "--port", "8000"]
Requirements
- Python 3.9+
- PostgreSQL 14+
- ~512 MB RAM minimum
- Anthropic API key (for AI fix generation)
- GitHub OAuth App (for repository integration)
Database
The schema is managed with Alembic (alembic.ini, migrations/versions). Railway runs alembic upgrade head before every deploy; a failed migration aborts the deploy and the running version keeps serving. The app and worker also upgrade on startup, which is a no-op once the schema is current.
alembic upgrade head # apply pending migrations (DATABASE_URL must be set)
alembic revision -m "describe" # start a new migration; write it as raw SQL with op.execute
Settings
| Variable | Default | Purpose |
|---|---|---|
DB_POOL_MAX_CONNECTIONS |
10 | Connections per process (web and worker each have one pool) |
DB_POOL_MIN_CONNECTIONS |
1 | Connections opened at startup |
DB_STATEMENT_TIMEOUT_MS |
30000 | App queries running longer are cancelled; 0 disables |
Storage
Repeats of a crash carry near identical payloads, so only the latest 20 events of each exception group keep their full payload (EVENT_PAYLOADS_KEPT_PER_GROUP in commons/autofix_config.py). Older events keep a summary (type, message, file, line) and still count toward the group. Repeats the SDK folds into one deduplicated_pattern event are added to the group's count without storing their sample. Each ingest batch is one transaction.
Key Tables
| Table | Purpose |
|---|---|
organizations |
Org settings, tier, Stripe billing |
events |
Exception events from the SDK, each linked to its group |
exception_groups |
Deduplicated exception groups with counts and GitHub links |
autofix_state |
Autofix queue: state, stage, attempts and history per group |
fix_attempts |
AI fix generation attempts |
llmcoder_messages |
Claude Code SDK conversation logs |
fix_suggestions |
Generated fix code (no-repo flow) |
usage_tracking |
Event/fix count audit trail |
Backup
Back up the PostgreSQL database regularly. The events and fix_attempts tables contain the most critical data.