The production portfolio: systems we build, run, and stay calm about.
Everything below was designed, built, and is operated today by the principal engineer of Capybara Consulting — not staged demos, not team projects observed from the sidelines. Production details are available under NDA on a call, and any subsystem can be demonstrated in a clean reference build on request.
Production Systems
Production LLM Platform
● LIVE — operational, scheduled twice-daily production runs
- Complete custom-LLM pipeline driving a live event-classification and calibration system
- Data curation → LoRA/DPO fine-tuning (Axolotl, TRL, Unsloth) → quantized serving (vLLM, llama.cpp/GGUF)
- Structured-output enforcement (Outlines + Pydantic schemas) — schema-guaranteed JSON for downstream systems, not "usually valid"
- Self-managed GPU capacity across multiple providers: automated VM lifecycle, GPU health checks, ephemeral checkpointing resilient to spot eviction
- Evaluation harnesses and experiment tracking (Weights & Biases); quantization tradeoff tuning for latency vs quality
Stack: PyTorch · Axolotl · TRL · Unsloth · vLLM · llama.cpp · Outlines · W&B · multi-cloud GPU
Real-Time Analytics & Decision Platform
● LIVE — full-stack, zero-downtime through node and storage failures
- FastAPI + PostgreSQL + nginx on multi-node Kubernetes; live data ingestion; calibration models; deploy-time decisions with confidence thresholds
- Decision accuracy: 92.9% out-of-sample vs a 61.8% baseline
- Full historical replay for backtesting, engineered against lookahead bias
- Scheduled pipelines, alerting, dashboards — zero-downtime through every infrastructure incident this quarter
Stack: FastAPI · PostgreSQL · nginx · k3s · Prometheus · Grafana · Python data stack
Distributed Kubernetes + Ceph Storage Platform
● LIVE — multi-node production cluster
- k3s + Rook/Ceph (CephFS + RBD) with Prometheus/Grafana/Alertmanager observability
- Automated node crash-loop recovery and boot-order gating for storage-dependent workloads
- CephFS mount-lifecycle deadlocks and Rook monitor failover loops diagnosed and fixed in production — runbooks written so they stay fixed
Stack: k3s · Rook · Ceph · systemd · Prometheus · Grafana · Alertmanager
Security-First Network Architecture
● LIVE — multi-site zero-trust operation
- IPv6 overlay networking, WireGuard mesh, DNS privacy layer (Unbound + DoT), Tor-based egress for sensitive workloads
- Hardened self-hosted messaging infrastructure; multi-layer firewalling, minimal attack surface
- Threat-model-first design, not compliance-checkbox design
Stack: IPv6 · WireGuard · Unbound · Tor · firewalling · OPSEC
Self-Hosted Digital-Asset Infrastructure
● LIVE — production multisig operations
- Solidity deployment and verification workflows; self-hosted payment processing (BTCPay Server)
- Multisig treasury operations tooling: key ceremonies, transaction staging, hardware-wallet signing workflows
- Full Safe{Wallet} infrastructure self-hosted on Kubernetes — transaction indexer, config service, client gateway, web UI — production multisig operations with zero third-party custody, operated for a private family-office client under NDA
Stack: Solidity · Gnosis Safe · Safe{Wallet} stack · BTCPay · hardware-wallet ops