Server Configuration

ThinkingSDK server settings control how exceptions are processed, how the autofix pipeline operates, and how billing works.


Environment Variables

The server reads configuration from environment variables. Set these in your .env file or deployment environment.

Core Settings

# Database
DATABASE_URL=postgresql://user:password@host:5432/thinkingsdk

# Server
APP_URL=https://your-domain.com
PORT=8000

# Authentication
OAUTH_CLIENT_ID=your_google_oauth_client_id
OAUTH_CLIENT_SECRET=your_google_oauth_secret
SESSION_SECRET=your_random_session_secret

GitHub Integration

# GitHub OAuth App credentials (for connecting user repos)
GITHUB_CLIENT_ID=your_github_oauth_client_id
GITHUB_CLIENT_SECRET=your_github_oauth_secret

When a user connects their GitHub account, ThinkingSDK can:

AI / LLM Configuration

# Anthropic (Claude) — primary fix generation engine
ANTHROPIC_API_KEY=sk-ant-your-key

# OpenAI (fallback)
OPENAI_API_KEY=sk-your-key

ThinkingSDK uses Claude as the primary engine for:

Stripe Billing

STRIPE_SECRET_KEY=sk_test_your_key
STRIPE_PUBLISHABLE_KEY=pk_test_your_key
STRIPE_WEBHOOK_SECRET=whsec_your_secret

Autofix Pipeline

The autofix pipeline turns captured exceptions into fixes. A worker claims one exception group at a time and runs it through a short list of steps.

Pipeline Phases

Phase What It Does
Phase 1: Intake Groups exceptions, checks org tier and usage limits, queues new groups
Phase 2: Setup Creates GitHub issue, clones repository, sets up workspace
Phase 3: Fix Generation Analyzes code, generates fix, evaluates with judge, runs tests
Phase 4: PR Creation Creates pull request, updates issue with PR link

Pipeline Behavior by Tier

Tier Crash Capture Time Capsule Dashboard Autofix + PR
Free Yes Yes Yes No
Contact Founders Yes Yes Yes Yes

Free tier users get full crash capture, Time Capsule viewing, and dashboard access. The autofix pipeline (AI fix generation + PR creation) requires a paid subscription.

States and Stages

Each exception group has a state (what the worker does with it) and a stage (how far it got):

queued -> running -> succeeded
                  -> queued     retryable error, retried with exponential backoff
                  -> waiting    needs you, e.g. connect GitHub, then press Retry Fix
                  -> failed     permanent error or retries used up

Stages, in order:

with a repository:    issue -> workspace -> fix -> pull_request
without a repository: analysis -> suggestion

A retry resumes after the last stage whose result is saved: an existing GitHub issue is never created twice and a group that already has a pull request is not fixed again.

Retries

Remote Client Configuration

Control SDK behavior server-side without redeploying client applications.

API Endpoints

GET  /api/v1/client-config          — Fetch current config
PUT  /api/v1/client-config          — Update config
GET  /api/v1/client-config/history  — View change history
DELETE /api/v1/client-config        — Reset to defaults

Example: Change Sample Rate Remotely

curl -X PUT https://api.thinkingsdk.ai/api/v1/client-config \
  -H "X-THINKINGSDK-KEY: tsdk_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "instrumentation": {
      "sample_rate": 0.1
    }
  }'

All connected clients with remote_config_enabled: true will pick up this change within their poll interval (default: 5 minutes).

Configurable Remotely


API Reference

Event Ingestion

POST /ingest

Accepts exception events from the SDK client. Supports batch and single events, gzip compression.

Headers (as sent by the Python client 0.1.6):
- X-Session-Token: sess_...: session from POST /auth/session; the client falls back to X-THINKINGSDK-KEY: tsdk_your_api_key when it cannot get one
- Content-Type: application/json
- Content-Encoding: gzip: bodies over 1 KiB
- User-Agent: ThinkingSDK-Python/0.1.6 (Python 3.12.11; Darwin)
- X-SDK-Language: Python
- X-SDK-Version: 0.1.6

Repeats of a crash within the client's 15 minute deduplication window arrive as one deduplicated_pattern event; its data.frequency is added to the group's occurrence count.

Response:

{
  "status": "ok",
  "events_received": 5,
  "events_stored": 5
}

Exception Flow

GET /api/event-flow/{event_id}

Returns the complete processing flow for an exception: event details, exception group, state transitions, fix attempts, and fix suggestions.

Crash Context (Time Capsule)

GET /api/crash-context/{event_id}

Returns the full frozen crash state: stack frames with source code, local/global variables at every frame, breadcrumbs, system snapshot, and AI fix data.

Billing Usage

GET /api/billing/usage

Returns current event count, fix attempt count, and tier limits.

Manual Retry

POST /api/manual-retry/{exception_group_id}

Manually trigger a retry for a failed fix attempt. Resets the iteration counter and starts a fresh fix attempt.


Deployment

ThinkingSDK server runs on any platform that supports Python and PostgreSQL.

# Procfile
web: uvicorn webapp.server:app --host 0.0.0.0 --port $PORT

Docker

FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
CMD ["uvicorn", "webapp.server:app", "--host", "0.0.0.0", "--port", "8000"]

Requirements


Database

The schema is managed with Alembic (alembic.ini, migrations/versions). Railway runs alembic upgrade head before every deploy; a failed migration aborts the deploy and the running version keeps serving. The app and worker also upgrade on startup, which is a no-op once the schema is current.

alembic upgrade head            # apply pending migrations (DATABASE_URL must be set)
alembic revision -m "describe"  # start a new migration; write it as raw SQL with op.execute

Settings

Variable Default Purpose
DB_POOL_MAX_CONNECTIONS 10 Connections per process (web and worker each have one pool)
DB_POOL_MIN_CONNECTIONS 1 Connections opened at startup
DB_STATEMENT_TIMEOUT_MS 30000 App queries running longer are cancelled; 0 disables

Storage

Repeats of a crash carry near identical payloads, so only the latest 20 events of each exception group keep their full payload (EVENT_PAYLOADS_KEPT_PER_GROUP in commons/autofix_config.py). Older events keep a summary (type, message, file, line) and still count toward the group. Repeats the SDK folds into one deduplicated_pattern event are added to the group's count without storing their sample. Each ingest batch is one transaction.

Key Tables

Table Purpose
organizations Org settings, tier, Stripe billing
events Exception events from the SDK, each linked to its group
exception_groups Deduplicated exception groups with counts and GitHub links
autofix_state Autofix queue: state, stage, attempts and history per group
fix_attempts AI fix generation attempts
llmcoder_messages Claude Code SDK conversation logs
fix_suggestions Generated fix code (no-repo flow)
usage_tracking Event/fix count audit trail

Backup

Back up the PostgreSQL database regularly. The events and fix_attempts tables contain the most critical data.