Capybara Consulting

AI · Security · Infrastructure

The production portfolio: systems we build, run, and stay calm about.

Everything below was designed, built, and is operated today by the principal engineer of Capybara Consulting — not staged demos, not team projects observed from the sidelines. Production details are available under NDA on a call, and any subsystem can be demonstrated in a clean reference build on request.

Production Systems

Production LLM Platform

● LIVE — operational, scheduled twice-daily production runs
  • Complete custom-LLM pipeline driving a live event-classification and calibration system
  • Data curation → LoRA/DPO fine-tuning (Axolotl, TRL, Unsloth) → quantized serving (vLLM, llama.cpp/GGUF)
  • Structured-output enforcement (Outlines + Pydantic schemas) — schema-guaranteed JSON for downstream systems, not "usually valid"
  • Self-managed GPU capacity across multiple providers: automated VM lifecycle, GPU health checks, ephemeral checkpointing resilient to spot eviction
  • Evaluation harnesses and experiment tracking (Weights & Biases); quantization tradeoff tuning for latency vs quality
Stack: PyTorch · Axolotl · TRL · Unsloth · vLLM · llama.cpp · Outlines · W&B · multi-cloud GPU

Real-Time Analytics & Decision Platform

● LIVE — full-stack, zero-downtime through node and storage failures
  • FastAPI + PostgreSQL + nginx on multi-node Kubernetes; live data ingestion; calibration models; deploy-time decisions with confidence thresholds
  • Decision accuracy: 92.9% out-of-sample vs a 61.8% baseline
  • Full historical replay for backtesting, engineered against lookahead bias
  • Scheduled pipelines, alerting, dashboards — zero-downtime through every infrastructure incident this quarter
Stack: FastAPI · PostgreSQL · nginx · k3s · Prometheus · Grafana · Python data stack

Distributed Kubernetes + Ceph Storage Platform

● LIVE — multi-node production cluster
  • k3s + Rook/Ceph (CephFS + RBD) with Prometheus/Grafana/Alertmanager observability
  • Automated node crash-loop recovery and boot-order gating for storage-dependent workloads
  • CephFS mount-lifecycle deadlocks and Rook monitor failover loops diagnosed and fixed in production — runbooks written so they stay fixed
Stack: k3s · Rook · Ceph · systemd · Prometheus · Grafana · Alertmanager

Security-First Network Architecture

● LIVE — multi-site zero-trust operation
  • IPv6 overlay networking, WireGuard mesh, DNS privacy layer (Unbound + DoT), Tor-based egress for sensitive workloads
  • Hardened self-hosted messaging infrastructure; multi-layer firewalling, minimal attack surface
  • Threat-model-first design, not compliance-checkbox design
Stack: IPv6 · WireGuard · Unbound · Tor · firewalling · OPSEC

Self-Hosted Digital-Asset Infrastructure

● LIVE — production multisig operations
  • Solidity deployment and verification workflows; self-hosted payment processing (BTCPay Server)
  • Multisig treasury operations tooling: key ceremonies, transaction staging, hardware-wallet signing workflows
  • Full Safe{Wallet} infrastructure self-hosted on Kubernetes — transaction indexer, config service, client gateway, web UI — production multisig operations with zero third-party custody, operated for a private family-office client under NDA
Stack: Solidity · Gnosis Safe · Safe{Wallet} stack · BTCPay · hardware-wallet ops

What This Means For You

Production AI systems need people who understand both the ML and the infrastructure it dies on. We've shipped the full vertical — kernel-level networking through model fine-tuning — and we operate everything we build. Fractional engagements (10–20 hrs/week) available through the A.Team network or directly.

Contact

Email: hello@capybaraconsulting.agency
XMPP: capybara-consulting@5222.de — OMEMO welcome.