Client SDK Configuration
The ThinkingSDK Python client (pip install thinkingsdk) is highly configurable. Configuration follows a priority order:
Function arguments → Environment variables → Config dict → YAML file → Defaults
Basic Setup
import thinkingsdk as thinking
thinking.start(
api_key="sk_live_your_key",
server_url="https://api.thinkingsdk.ai"
)
server_url defaults to http://localhost:8000, so always set it for the hosted service. If you leave out api_key, the SDK reads it from the THINKINGSDK_API_KEY environment variable.
Advanced Setup
thinking.start(
api_key="sk_live_your_key",
server_url="https://api.thinkingsdk.ai",
config={
'git_repositories': ['https://github.com/your-org/your-repo'],
'instrumentation': {
'capture_caught_exceptions': True,
'ignore_patterns': [r'migrations/'],
},
'sender': {
'batch_size': 100,
'retry_attempts': 5,
},
'privacy': {
'sanitize_data': True,
'custom_keys': ['internal_id'],
}
},
enable_logging=True
)
Configuration Reference
Options live under the section shown in their name (instrumentation, sender, queue, deduplication, privacy), in the config dict and in the YAML file alike. git_repositories is top level. A key at the wrong level is ignored.
Autofix Repository
| Option | Type | Default | Description |
|---|---|---|---|
git_repositories |
list | [] | GitHub repository URLs: https://github.com/owner/repo or git@github.com:owner/repo.git. The autofix pipeline uses the first entry to choose the repository flow |
Exception events carry a repository_context with this repository plus the branch and commit of the running checkout, read from the .git directory that contains the process's working directory (git itself is not invoked). When no branch can be read, as in most containers or on a detached HEAD, the branch falls back to main; outside a checkout the commit is empty. Without git_repositories, events are sent with autofix disabled. Worker thread crashes do not carry repository context.
Instrumentation
Controls what data the SDK captures from your application. Unhandled exceptions are always captured through sys.excepthook (main thread) and threading.excepthook (worker threads). The options below other than capture_caught_exceptions configure the caught exception tracer and have no effect while it is off.
| Option | Type | Default | Description |
|---|---|---|---|
instrumentation.capture_caught_exceptions |
bool | False | Also report exceptions that your code or your web framework catches (see Caught Exceptions) |
instrumentation.max_locals |
int | 5 | Variables kept in the raising frame's locals and globals summary |
instrumentation.max_local_length |
int | 120 | Maximum string length for those values |
instrumentation.capture_memory |
bool | False | Add process memory usage to each event |
instrumentation.capture_source_lines |
bool | False | Add the current source line to each event |
instrumentation.exceptions_only |
bool | True | Send only exception events from the tracer (recommended) |
Caught Exceptions
Since 0.1.5 caught exceptions are not reported by default. That includes request exceptions in FastAPI, Flask and Django: the framework catches them and returns a 500, so they never reach sys.excepthook. To report them, opt in:
thinking.start(
api_key="sk_live_your_key",
server_url="https://api.thinkingsdk.ai",
config={'instrumentation': {'capture_caught_exceptions': True}}
)
or set THINKINGSDK_CAPTURE_CAUGHT_EXCEPTIONS=true. The tracer runs on every raise in the process, including ordinary control flow such as cache misses and iterator exhaustion, so measure the overhead on your own workload first. It uses sys.monitoring on Python 3.12+ and falls back to the much slower sys.settrace on older versions. An exception that propagates through several of your functions is reported once per frame it passes through. Tracer events also carry the local and global variables of every traceback frame (up to 50 per frame) and the source lines around the error; crash reports from sys.excepthook alone do not.
Sampling & Filtering
Unhandled exceptions are never sampled or filtered and the tracer's default strategic sampler always keeps exception events, so instrumentation.sample_rate (default 1.0, THINKINGSDK_SAMPLE_RATE) does not reduce exception volume. Repeats are aggregated by deduplication instead. The tracer can skip code with:
| Option | Type | Default | Description |
|---|---|---|---|
instrumentation.ignore_patterns |
list | [] | Regex patterns matched against file paths to skip |
instrumentation.ignore_functions |
list | [] | Extra function names to skip (dunder methods such as __init__ are always skipped) |
Network & Batching
Controls how events are sent to ThinkingSDK servers.
| Option | Type | Default | Description |
|---|---|---|---|
sender.batch_size |
int | 50 | Events per batch |
sender.max_batch_wait |
float | 2.0 | Seconds to wait before sending partial batch |
sender.max_batch_bytes |
int | 1048576 (1 MiB) | Payload budget per batch |
sender.retry_attempts |
int | 3 | Retry count on failure |
sender.backoff_factor |
float | 1.0 | Exponential backoff multiplier |
sender.circuit_breaker_threshold |
int | 5 | Consecutive failures before circuit opens |
sender.circuit_breaker_timeout |
int | 60 | Seconds before retrying after circuit opens |
sender.request_timeout |
int | 10 | HTTP request timeout in seconds |
Batches over 1 KB are gzip compressed. After a failed batch the sender waits with exponential backoff, capped at 60 seconds, before sending the next one.
Event Queue
| Option | Type | Default | Description |
|---|---|---|---|
queue.maxsize |
int | 10000 | Maximum events to buffer |
queue.max_bytes |
int | 8388608 (8 MiB) | Maximum serialized bytes to buffer |
queue.max_event_bytes |
int | 262144 (256 KiB) | Maximum size of one report; larger reports are dropped |
queue.drop_strategy |
string | "oldest" | What to drop on overflow: "oldest" or "newest" |
Before queueing, each report is bounded: long strings are truncated and on large reports breadcrumbs and context are cut before the exception and repository context.
Deduplication
The first occurrence of a crash is sent immediately. Repeats with the same exception type and stack location are aggregated into one sample and a count, sent as a deduplicated_pattern event (the count is in data.frequency) once the window closes. Custom events are not deduplicated.
| Option | Type | Default | Description |
|---|---|---|---|
deduplication.window_size_ms |
int | 900000 (15 min) | Aggregation window, measured from the first occurrence |
deduplication.flush_interval_ms |
int | 900000 (15 min) | The summary is sent after the shorter of this and window_size_ms |
deduplication.max_patterns |
int | 1000 | Distinct crash signatures tracked; the oldest is evicted |
deduplication.max_bytes |
int | 2097152 (2 MiB) | Total bytes of retained samples |
deduplication.max_sample_bytes |
int | 65536 (64 KiB) | Maximum size of one retained sample; larger reports skip deduplication |
Repeat counts still pending when the process exits are not sent.
Privacy & Security
| Option | Type | Default | Description |
|---|---|---|---|
privacy.sanitize_data |
bool | True | Enable PII scrubbing |
privacy.custom_keys |
list | [] | Extra key names to redact (lowercase, matched as substrings) |
privacy.excluded_keys |
list | [] | Key names never redacted |
privacy.custom_patterns |
dict | {} | Extra regexes to redact, as {name: pattern} |
privacy.scrub_emails |
bool | True | Partially mask email addresses |
privacy.scrub_ips |
bool | False | Redact IP addresses |
privacy.scrub_uuids |
bool | False | Redact UUIDs |
privacy.replacement |
string | "[REDACTED]" | Replacement text for pattern matches |
With scrubbing on, values under keys containing password, secret, token, api_key, auth, session, cookie, ssn, credit_card and similar names are redacted, as are values that look like card numbers (last 4 digits kept), SSNs, phone numbers, JWTs, AWS access keys, AWS secret keys (a standalone 40 character key mixing upper case, lower case and digits, so git SHAs and file paths are kept) and GitHub or Stripe keys. Scrubbing runs on exception events; data passed to track_event, track_metric, mark_feature_usage and timer is sent as given.
Environment Variables
These variables override the matching config dict and YAML values. Booleans accept true, 1, yes or on.
THINKINGSDK_API_KEY=sk_live_your_key # used when start() gets no api_key
THINKINGSDK_CAPTURE_CAUGHT_EXCEPTIONS=true # instrumentation.capture_caught_exceptions
THINKINGSDK_SAMPLE_RATE=0.5 # instrumentation.sample_rate
THINKINGSDK_BATCH_SIZE=100 # sender.batch_size
THINKINGSDK_QUEUE_SIZE=20000 # queue.maxsize
THINKINGSDK_ENABLE_LOGGING=true
THINKINGSDK_LOG_LEVEL=DEBUG
THINKINGSDK_DISABLE_AUTO=true # start() does nothing (any non-empty value)
THINKINGSDK_SERVER_URL is not read directly: reference it from thinkingsdk.yaml as ${THINKINGSDK_SERVER_URL}. If python-dotenv is installed, a .env file in the working directory or a parent directory is loaded on import.
YAML Configuration File
For complex configurations, use a YAML file:
thinking.start(config_file="thinkingsdk.yaml")
Without config_file, the SDK looks for thinkingsdk.yaml in the working directory and its parents. api_key and server_url arguments override the file. ${VAR} or ${VAR:-default} in the file expands from the environment.
Example thinkingsdk.yaml:
api_key_source: "env:THINKINGSDK_API_KEY" # or "file:~/.thinkingsdk/key" or "keyring:<service>"
server_url: "${THINKINGSDK_SERVER_URL:-https://api.thinkingsdk.ai}"
enabled: true # false: start() does nothing
debug: false # true: enable SDK logging
git_repositories:
- "https://github.com/your-org/your-repo"
instrumentation:
capture_caught_exceptions: false
ignore_patterns:
- "migrations/"
sender:
batch_size: 100
retry_attempts: 5
privacy:
sanitize_data: true
custom_keys:
- internal_id
keyring: needs pip install "thinkingsdk[keyring]". Sections from older sample files (tracking, export, performance, top level strategic_sampling) have no effect.
Breadcrumbs
Breadcrumbs track the execution trail leading up to a crash. They're automatically captured from integrations, or you can add them manually.
Automatic Breadcrumbs
The SDK auto-captures breadcrumbs from:
- HTTP requests: outgoing calls through
http.client(includingurllibandrequests) with method, URL and status code; incoming Flask and Django requests - Database queries: SQLAlchemy and psycopg2 (parameters not captured), PyMongo (queries sanitized)
- Cache operations: Redis commands
- Logging:
loggingrecords at INFO and above that pass your loggers' levels - Print statements: console output
- Subprocess calls: commands started through
subprocess
Manual Breadcrumbs
thinking.add_breadcrumb(
message="User clicked checkout",
category="user",
level="info",
data={"cart_total": 99.99, "items": 3}
)
Categories: any string (default default). Automatic breadcrumbs use http, db, cache, console, subprocess and custom (from track_event); log records use the logger name.
Levels: debug, info, warning, error (fatal for critical log records)
Up to 100 breadcrumbs are retained; the 50 most recent are attached to each exception event. Worker thread crash events carry none.
Custom Events & Metrics
Track custom application events alongside crash data:
# Track a business event
thinking.track_event(
"payment_processed",
data={"amount": 99.99, "currency": "USD"},
level="info"
)
# Track a performance metric
thinking.track_metric(
"api_latency",
value=245.5,
unit="ms",
tags={"endpoint": "/api/users"}
)
# Mark feature usage
thinking.mark_feature_usage(
"dark_mode",
metadata={"enabled_by": "user_preference"}
)
# Timer context manager
with thinking.timer("database_query", tags={"table": "users"}):
result = db.query("SELECT * FROM users")
Custom events skip deduplication and PII scrubbing. These functions and add_breadcrumb raise RuntimeError if called before start().
Context Management
Set context that's attached to custom events and caught exception events:
# Scope context to a block (works with `with` and `async with`)
with thinking.context(user_id="usr_123", request_id="req_abc"):
process_request()
# Or set global context
thinking.set_context({"user_id": "usr_123"})
thinking.add_context("request_id", "req_abc")
thinking.clear_context()
Crash reports from sys.excepthook do not include this context.
Lifecycle
# Start the SDK (raises RuntimeError if it is already running)
thinking.start(api_key="sk_live_your_key", server_url="https://api.thinkingsdk.ai")
# Check if active
if thinking.is_active():
print("SDK is running")
# Get statistics
stats = thinking.get_stats()
print(f"Events sent: {stats['sender']['total_sent']}")
print(f"Events dropped: {stats['queue']['dropped_count']}")
# Graceful shutdown: restores the original hooks and waits up to 5 seconds for a final send
thinking.stop(timeout=5.0)
An exit handler calls stop() automatically, so a crash that ends the process is still sent. Delivery is best effort: exhausted retries, buffer pressure and process exit can lose events.
Framework Integrations
The SDK automatically integrates with detected frameworks:
| Framework | Auto-Detected | Captures |
|---|---|---|
| Flask | Yes | Request breadcrumbs, app config on exception events |
| FastAPI | Yes | FastAPI and Starlette versions plus async state on exception events (no request breadcrumbs) |
| Django | Yes | Request breadcrumbs, middleware and safe settings on exception events |
| SQLAlchemy | Yes | Query breadcrumbs (parameters not captured) |
| psycopg2 | Yes | PostgreSQL query breadcrumbs (parameters not captured) |
| PyMongo | Yes | MongoDB query breadcrumbs (sanitized) |
| Redis | Yes | Command breadcrumbs (values not captured) |
| Python logging | Yes | Log records as breadcrumbs |
No configuration needed: the SDK detects these libraries when start() runs. Integrations add breadcrumbs and context; they do not report the request exceptions a framework catches (see Caught Exceptions).
Production Characteristics
- Idle until a crash: the excepthooks add no per call cost; the caught exception tracer is opt in
- Bounded buffers: 8 MiB of queued reports and 2 MiB of deduplication samples by default (payload limits, not a cap on process memory)
- Non-blocking: All network I/O happens in a background thread
- Circuit breaker: Pauses sending for 60 seconds after 5 consecutive failures
- Graceful degradation: SDK failures never crash your application
- Thread-safe: All components are thread-safe
- gzip compression: Batches over 1 KB are compressed before sending