Skip to main content
EVOKORE// BROWSE
>

./browse/prompts

13 NODES
šŸ“text•6 months ago

Evaluation and Regression Gate

Create evaluation runs, submit results, compute metrics, save baselines, and triage regressions before shipping changes.

analysis
⭐1
# Evaluation and Regression Gate Imported from curated first-party documentation sources. ## What this covers Use this workflow to formalize benchmarking, regression detection, and baseline management around agent or product changes. ## Use this when - Validating changes against prior baselines - Making regressions visible before rollout - Capturing repeatable evaluation evidence ## Expected outcomes - Evaluation runs produce comparable metrics - Baselines are stored after successful validation - Regression triage becomes a workflow instead of an afterthought ## Source synthesis - AGENT33/docs/functionality-and-workflows.md (https://github.com/mattmre/AGENT33/blob/main/docs/functionality-and-workflows.md) - AGENT33/docs/use-cases.md (https://github.com/mattmre/AGENT33/blob/main/docs/use-cases.md) ## Dedupe notes Combines AGENT33 evaluation lifecycle definitions with evaluation and regression use-case framing. ## Source excerpts ### AGENT33/docs/functionality-and-workflows.md ### 4.3 Evaluation Lifecycle Flow: 1. Create run (`/v1/evaluations/runs`) 2. Submit task results (`/runs/{id}/results`) 3. Compute metrics + gate report 4. Save baseline (`/runs/{id}/baseline`) 5. Triage/resolve regressions ### AGENT33/docs/use-cases.md ## 4. Evaluation and Regression Gates Goal: - Quantify quality and block regressions across PR/merge/release gates. Use these modules: - `api/routes/evaluations.py` - `evaluation/service.py` - `evaluation/gates.py` - `evaluation/regression.py` Typical flow: 1. Create run for gate type (`G-PR`, `G-MRG`, `G-REL`, `G-MON`). 2. Submit task results and quality metadata. 3. Compute metrics and gate verdict. 4. Save baseline for future comparison. 5. Triage/resolve regression records. Best fit: - Teams with golden-task style quality gates.
šŸ‘0
šŸ‘ļø0
docs
šŸ¤–system prompt•7 months ago

architecture-decision-records

Write and maintain Architecture Decision Records (ADRs) following

coding
⭐1
# Architecture Decision Records Comprehensive patterns for creating, maintaining, and managing Architecture Decision Records (ADRs) that capture the context and rationale behind significant technical decisions. ## When to Use This Skill - Making significant architectural decisions - Documenting technology choices - Recording design trade-offs - Onboarding new team members - Reviewing historical decisions - Establishing decision-making processes ## Core Concepts ### 1. What is an ADR? An Architecture Decision Record captures: - **Context**: Why we needed to make a decision - **Decision**: What we decided - **Consequences**: What happens as a result ### 2. When to Write an ADR | Write ADR | Skip ADR | | -------------------------- | ---------------------- | | New framework adoption | Minor version upgrades | | Database technology choice | Bug fixes | | API design patterns | Implementation details | | Security architecture | Routine maintenance | | Integration patterns | Configuration changes | ### 3. ADR Lifecycle ``` Proposed → Accepted → Deprecated → Superseded ↓ Rejected ``` ## Templates ### Template 1: Standard ADR (MADR Format) ```markdown # ADR-0001: Use PostgreSQL as Primary Database ## Status Accepted ## Context We need to select a primary database for our new e-commerce platform. The system will handle: - ~10,000 concurrent users - Complex product catalog with hierarchical categories - Transaction processing for orders and payments - Full-text search for products - Geospatial queries for store locator The team has experience with MySQL, PostgreSQL, and MongoDB. We need ACID compliance for financial transactions. ## Decision Drivers - **Must have ACID compliance** for payment processing - **Must support complex queries** for reporting - **Should support full-text search** to reduce infrastructure complexity - **Should have good JSON support** for flexible product attributes - **Team familiarity** reduces onboarding time ## Considered Options ### Option 1: PostgreSQL - **Pros**: ACID compliant, excellent JSON support (JSONB), built-in full-text search, PostGIS for geospatial, team has experience - **Cons**: Slightly more complex replication setup than MySQL ### Option 2: MySQL - **Pros**: Very familiar to team, simple replication, large community - **Cons**: Weaker JSON support, no built-in full-text search (need Elasticsearch), no geospatial without extensions ### Option 3: MongoDB - **Pros**: Flexible schema, native JSON, horizontal scaling - **Cons**: No ACID for multi-document transactions (at decision time), team has limited experience, requires schema design discipline ## Decision We will use **PostgreSQL 15** as our primary database. ## Rationale PostgreSQL provides the best balance of: 1. **ACID compliance** essential for e-commerce transactions 2. **Built-in capabilities** (full-text search, JSONB, PostGIS) reduce infrastructure complexity 3. **Team familiarity** with SQL databases reduces learning curve 4. **Mature ecosystem** with excellent tooling and community support The slight complexity in replication is outweighed by the reduction in additional services (no separate Elasticsearch needed). ## Consequences ### Positive - Single database handles transactions, search, and geospatial queries - Reduced operational complexity (fewer services to manage) - Strong consistency guarantees for financial data - Team can leverage existing SQL expertise ### Negative - Need to learn PostgreSQL-specific features (JSONB, full-text search syntax) - Vertical scaling limits may require read replicas sooner - Some team members need PostgreSQL-specific training ### Risks - Full-text search may not scale as well as dedicated search engines - Mitigation: Design for potential Elasticsearch addition if needed ## Implementation Notes - Use JSONB for flexible product attributes - Implement connection pooling with PgBouncer - Set up streaming replication for read replicas - Use pg_trgm extension for fuzzy search ## Related Decisions - ADR-0002: Caching Strategy (Redis) - complements database choice - ADR-0005: Search Architecture - may supersede if Elasticsearch needed ## References - [PostgreSQL JSON Documentation](https://www.postgresql.org/docs/current/datatype-json.html) - [PostgreSQL Full Text Search](https://www.postgresql.org/docs/current/textsearch.html) - Internal: Performance benchmarks in `/docs/benchmarks/database-comparison.md` ``` ### Template 2: Lightweight ADR ```markdown # ADR-0012: Adopt TypeScript for Frontend Development **Status**: Accepted **Date**: 2024-01-15 **Deciders**: @alice, @bob, @charlie ## Context Our React codebase has grown to 50+ components with increasing bug reports related to prop type mismatches and undefined errors. PropTypes provide runtime-only checking. ## Decision Adopt TypeScript for all new frontend code. Migrate existing code incrementally. ## Consequences **Good**: Catch type errors at compile time, better IDE support, self-documenting code. **Bad**: Learning curve for team, initial slowdown, build complexity increase. **Mitigations**: TypeScript training sessions, allow gradual adoption with `allowJs: true`. ``` ### Template 3: Y-Statement Format ```markdown # ADR-0015: API Gateway Selection In the context of **building a microservices architecture**, facing **the need for centralized API management, authentication, and rate limiting**, we decided for **Kong Gateway** and against **AWS API Gateway and custom Nginx solution**, to achieve **vendor independence, plugin extensibility, and team familiarity with Lua**, accepting that **we need to manage Kong infrastructure ourselves**. ``` ### Template 4: ADR for Deprecation ```markdown # ADR-0020: Deprecate MongoDB in Favor of PostgreSQL ## Status Accepted (Supersedes ADR-0003) ## Context ADR-0003 (2021) chose MongoDB for user profile storage due to schema flexibility needs. Since then: - MongoDB's multi-document transactions remain problematic for our use case - Our schema has stabilized and rarely changes - We now have PostgreSQL expertise from other services - Maintaining two databases increases operational burden ## Decision Deprecate MongoDB and migrate user profiles to PostgreSQL. ## Migration Plan 1. **Phase 1** (Week 1-2): Create PostgreSQL schema, dual-write enabled 2. **Phase 2** (Week 3-4): Backfill historical data, validate consistency 3. **Phase 3** (Week 5): Switch reads to PostgreSQL, monitor 4. **Phase 4** (Week 6): Remove MongoDB writes, decommission ## Consequences ### Positive - Single database technology reduces operational complexity - ACID transactions for user data - Team can focus PostgreSQL expertise ### Negative - Migration effort (~4 weeks) - Risk of data issues during migration - Lose some schema flexibility ## Lessons Learned Document from ADR-0003 experience: - Schema flexibility benefits were overestimated - Operational cost of multiple databases was underestimated - Consider long-term maintenance in technology decisions ``` ### Template 5: Request for Comments (RFC) Style ```markdown # RFC-0025: Adopt Event Sourcing for Order Management ## Summary Propose adopting event sourcing pattern for the order management domain to improve auditability, enable temporal queries, and support business analytics. ## Motivation Current challenges: 1. Audit requirements need complete order history 2. "What was the order state at time X?" queries are impossible 3. Analytics team needs event stream for real-time dashboards 4. Order state reconstruction for customer support is manual ## Detailed Design ### Event Store ``` OrderCreated { orderId, customerId, items[], timestamp } OrderItemAdded { orderId, item, timestamp } OrderItemRemoved { orderId, itemId, timestamp } PaymentReceived { orderId, amount, paymentId, timestamp } OrderShipped { orderId, trackingNumber, timestamp } ``` ### Projections - **CurrentOrderState**: Materialized view for queries - **OrderHistory**: Complete timeline for audit - **DailyOrderMetrics**: Analytics aggregation ### Technology - Event Store: EventStoreDB (purpose-built, handles projections) - Alternative considered: Kafka + custom projection service ## Drawbacks - Learning curve for team - Increased complexity vs. CRUD - Need to design events carefully (immutable once stored) - Storage growth (events never deleted) ## Alternatives 1. **Audit tables**: Simpler but doesn't enable temporal queries 2. **CDC from existing DB**: Complex, doesn't change data model 3. **Hybrid**: Event source only for order state changes ## Unresolved Questions - [ ] Event schema versioning strategy - [ ] Retention policy for events - [ ] Snapshot frequency for performance ## Implementation Plan 1. Prototype with single order type (2 weeks) 2. Team training on event sourcing (1 week) 3. Full implementation and migration (4 weeks) 4. Monitoring and optimization (ongoing) ## References - [Event Sourcing by Martin Fowler](https://martinfowler.com/eaaDev/EventSourcing.html) - [EventStoreDB Documentation](https://www.eventstore.com/docs) ``` ## ADR Management ### Directory Structure ``` docs/ ā”œā”€ā”€ adr/ │ ā”œā”€ā”€ README.md # Index and guidelines │ ā”œā”€ā”€ template.md # Team's ADR template │ ā”œā”€ā”€ 0001-use-postgresql.md │ ā”œā”€ā”€ 0002-caching-strategy.md │ ā”œā”€ā”€ 0003-mongodb-user-profiles.md # [DEPRECATED] │ └── 0020-deprecate-mongodb.md # Supersedes 0003 ``` ### ADR Index (README.md) ```markdown # Architecture Decision Records This directory contains Architecture Decision Records (ADRs) for [Project Name]. ## Index | ADR | Title | Status | Date | | ------------------------------------- | ---------------------------------- | ---------- | ---------- | | [0001](0001-use-postgresql.md) | Use PostgreSQL as Primary Database | Accepted | 2024-01-10 | | [0002](0002-caching-strategy.md) | Caching Strategy with Redis | Accepted | 2024-01-12 | | [0003](0003-mongodb-user-profiles.md) | MongoDB for User Profiles | Deprecated | 2023-06-15 | | [0020](0020-deprecate-mongodb.md) | Deprecate MongoDB | Accepted | 2024-01-15 | ## Creating a New ADR 1. Copy `template.md` to `NNNN-title-with-dashes.md` 2. Fill in the template 3. Submit PR for review 4. Update this index after approval ## ADR Status - **Proposed**: Under discussion - **Accepted**: Decision made, implementing - **Deprecated**: No longer relevant - **Superseded**: Replaced by another ADR - **Rejected**: Considered but not adopted ``` ### Automation (adr-tools) ```bash # Install adr-tools brew install adr-tools # Initialize ADR directory adr init docs/adr # Create new ADR adr new "Use PostgreSQL as Primary Database" # Supersede an ADR adr new -s 3 "Deprecate MongoDB in Favor of PostgreSQL" # Generate table of contents adr generate toc > docs/adr/README.md # Link related ADRs adr link 2 "Complements" 1 "Is complemented by" ``` ## Review Process ```markdown ## ADR Review Checklist ### Before Submission - [ ] Context clearly explains the problem - [ ] All viable options considered - [ ] Pros/cons balanced and honest - [ ] Consequences (positive and negative) documented - [ ] Related ADRs linked ### During Review - [ ] At least 2 senior engineers reviewed - [ ] Affected teams consulted - [ ] Security implications considered - [ ] Cost implications documented - [ ] Reversibility assessed ### After Acceptance - [ ] ADR index updated - [ ] Team notified - [ ] Implementation tickets created - [ ] Related documentation updated ``` ## Best Practices ### Do's - **Write ADRs early** - Before implementation starts - **Keep them short** - 1-2 pages maximum - **Be honest about trade-offs** - Include real cons - **Link related decisions** - Build decision graph - **Update status** - Deprecate when superseded ### Don'ts - **Don't change accepted ADRs** - Write new ones to supersede - **Don't skip context** - Future readers need background - **Don't hide failures** - Rejected decisions are valuable - **Don't be vague** - Specific decisions, specific consequences - **Don't forget implementation** - ADR without action is waste ## Resources - [Documenting Architecture Decisions (Michael Nygard)](https://cognitect.com/blog/2011/11/15/documenting-architecture-decisions) - [MADR Template](https://adr.github.io/madr/) - [ADR GitHub Organization](https://adr.github.io/) - [adr-tools](https://github.com/npryce/adr-tools)
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

k8s-security-policies

Implement Kubernetes security policies including NetworkPolicy,

architecture
⭐1
# Kubernetes Security Policies Comprehensive guide for implementing NetworkPolicy, PodSecurityPolicy, RBAC, and Pod Security Standards in Kubernetes. ## Purpose Implement defense-in-depth security for Kubernetes clusters using network policies, pod security standards, and RBAC. ## When to Use This Skill - Implement network segmentation - Configure pod security standards - Set up RBAC for least-privilege access - Create security policies for compliance - Implement admission control - Secure multi-tenant clusters ## Pod Security Standards ### 1. Privileged (Unrestricted) ```yaml apiVersion: v1 kind: Namespace metadata: name: privileged-ns labels: pod-security.kubernetes.io/enforce: privileged pod-security.kubernetes.io/audit: privileged pod-security.kubernetes.io/warn: privileged ``` ### 2. Baseline (Minimally restrictive) ```yaml apiVersion: v1 kind: Namespace metadata: name: baseline-ns labels: pod-security.kubernetes.io/enforce: baseline pod-security.kubernetes.io/audit: baseline pod-security.kubernetes.io/warn: baseline ``` ### 3. Restricted (Most restrictive) ```yaml apiVersion: v1 kind: Namespace metadata: name: restricted-ns labels: pod-security.kubernetes.io/enforce: restricted pod-security.kubernetes.io/audit: restricted pod-security.kubernetes.io/warn: restricted ``` ## Network Policies ### Default Deny All ```yaml apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: default-deny-all namespace: production spec: podSelector: {} policyTypes: - Ingress - Egress ``` ### Allow Frontend to Backend ```yaml apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: allow-frontend-to-backend namespace: production spec: podSelector: matchLabels: app: backend policyTypes: - Ingress ingress: - from: - podSelector: matchLabels: app: frontend ports: - protocol: TCP port: 8080 ``` ### Allow DNS ```yaml apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: allow-dns namespace: production spec: podSelector: {} policyTypes: - Egress egress: - to: - namespaceSelector: matchLabels: name: kube-system ports: - protocol: UDP port: 53 ``` **Reference:** See `assets/network-policy-template.yaml` ## RBAC Configuration ### Role (Namespace-scoped) ```yaml apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: pod-reader namespace: production rules: - apiGroups: [""] resources: ["pods"] verbs: ["get", "watch", "list"] ``` ### ClusterRole (Cluster-wide) ```yaml apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: secret-reader rules: - apiGroups: [""] resources: ["secrets"] verbs: ["get", "watch", "list"] ``` ### RoleBinding ```yaml apiVersion: rbac.authorization.k8s.io/v1 kind: RoleBinding metadata: name: read-pods namespace: production subjects: - kind: User name: jane apiGroup: rbac.authorization.k8s.io - kind: ServiceAccount name: default namespace: production roleRef: kind: Role name: pod-reader apiGroup: rbac.authorization.k8s.io ``` **Reference:** See `references/rbac-patterns.md` ## Pod Security Context ### Restricted Pod ```yaml apiVersion: v1 kind: Pod metadata: name: secure-pod spec: securityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000 seccompProfile: type: RuntimeDefault containers: - name: app image: myapp:1.0 securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: drop: - ALL ``` ## Policy Enforcement with OPA Gatekeeper ### ConstraintTemplate ```yaml apiVersion: templates.gatekeeper.sh/v1 kind: ConstraintTemplate metadata: name: k8srequiredlabels spec: crd: spec: names: kind: K8sRequiredLabels validation: openAPIV3Schema: type: object properties: labels: type: array items: type: string targets: - target: admission.k8s.gatekeeper.sh rego: | package k8srequiredlabels violation[{"msg": msg, "details": {"missing_labels": missing}}] { provided := {label | input.review.object.metadata.labels[label]} required := {label | label := input.parameters.labels[_]} missing := required - provided count(missing) > 0 msg := sprintf("missing required labels: %v", [missing]) } ``` ### Constraint ```yaml apiVersion: constraints.gatekeeper.sh/v1beta1 kind: K8sRequiredLabels metadata: name: require-app-label spec: match: kinds: - apiGroups: ["apps"] kinds: ["Deployment"] parameters: labels: ["app", "environment"] ``` ## Service Mesh Security (Istio) ### PeerAuthentication (mTLS) ```yaml apiVersion: security.istio.io/v1beta1 kind: PeerAuthentication metadata: name: default namespace: production spec: mtls: mode: STRICT ``` ### AuthorizationPolicy ```yaml apiVersion: security.istio.io/v1beta1 kind: AuthorizationPolicy metadata: name: allow-frontend namespace: production spec: selector: matchLabels: app: backend action: ALLOW rules: - from: - source: principals: ["cluster.local/ns/production/sa/frontend"] ``` ## Best Practices 1. **Implement Pod Security Standards** at namespace level 2. **Use Network Policies** for network segmentation 3. **Apply least-privilege RBAC** for all service accounts 4. **Enable admission control** (OPA Gatekeeper/Kyverno) 5. **Run containers as non-root** 6. **Use read-only root filesystem** 7. **Drop all capabilities** unless needed 8. **Implement resource quotas** and limit ranges 9. **Enable audit logging** for security events 10. **Regular security scanning** of images ## Compliance Frameworks ### CIS Kubernetes Benchmark - Use RBAC authorization - Enable audit logging - Use Pod Security Standards - Configure network policies - Implement secrets encryption at rest - Enable node authentication ### NIST Cybersecurity Framework - Implement defense in depth - Use network segmentation - Configure security monitoring - Implement access controls - Enable logging and monitoring ## Troubleshooting **NetworkPolicy not working:** ```bash # Check if CNI supports NetworkPolicy kubectl get nodes -o wide kubectl describe networkpolicy <name> ``` **RBAC permission denied:** ```bash # Check effective permissions kubectl auth can-i list pods --as system:serviceaccount:default:my-sa kubectl auth can-i '*' '*' --as system:serviceaccount:default:my-sa ``` ## Reference Files - `assets/network-policy-template.yaml` - Network policy examples - `assets/pod-security-template.yaml` - Pod security policies - `references/rbac-patterns.md` - RBAC configuration patterns ## Related Skills - `k8s-manifest-generator` - For creating secure manifests - `gitops-workflow` - For automated policy deployment
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

llm-evaluation

Implement comprehensive evaluation strategies for LLM applications

coding
⭐1
# LLM Evaluation Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing. ## When to Use This Skill - Measuring LLM application performance systematically - Comparing different models or prompts - Detecting performance regressions before deployment - Validating improvements from prompt changes - Building confidence in production systems - Establishing baselines and tracking progress over time - Debugging unexpected model behavior ## Core Evaluation Types ### 1. Automated Metrics Fast, repeatable, scalable evaluation using computed scores. **Text Generation:** - **BLEU**: N-gram overlap (translation) - **ROUGE**: Recall-oriented (summarization) - **METEOR**: Semantic similarity - **BERTScore**: Embedding-based similarity - **Perplexity**: Language model confidence **Classification:** - **Accuracy**: Percentage correct - **Precision/Recall/F1**: Class-specific performance - **Confusion Matrix**: Error patterns - **AUC-ROC**: Ranking quality **Retrieval (RAG):** - **MRR**: Mean Reciprocal Rank - **NDCG**: Normalized Discounted Cumulative Gain - **Precision@K**: Relevant in top K - **Recall@K**: Coverage in top K ### 2. Human Evaluation Manual assessment for quality aspects difficult to automate. **Dimensions:** - **Accuracy**: Factual correctness - **Coherence**: Logical flow - **Relevance**: Answers the question - **Fluency**: Natural language quality - **Safety**: No harmful content - **Helpfulness**: Useful to the user ### 3. LLM-as-Judge Use stronger LLMs to evaluate weaker model outputs. **Approaches:** - **Pointwise**: Score individual responses - **Pairwise**: Compare two responses - **Reference-based**: Compare to gold standard - **Reference-free**: Judge without ground truth ## Quick Start ```python from dataclasses import dataclass from typing import Callable import numpy as np @dataclass class Metric: name: str fn: Callable @staticmethod def accuracy(): return Metric("accuracy", calculate_accuracy) @staticmethod def bleu(): return Metric("bleu", calculate_bleu) @staticmethod def bertscore(): return Metric("bertscore", calculate_bertscore) @staticmethod def custom(name: str, fn: Callable): return Metric(name, fn) class EvaluationSuite: def __init__(self, metrics: list[Metric]): self.metrics = metrics async def evaluate(self, model, test_cases: list[dict]) -> dict: results = {m.name: [] for m in self.metrics} for test in test_cases: prediction = await model.predict(test["input"]) for metric in self.metrics: score = metric.fn( prediction=prediction, reference=test.get("expected"), context=test.get("context") ) results[metric.name].append(score) return { "metrics": {k: np.mean(v) for k, v in results.items()}, "raw_scores": results } # Usage suite = EvaluationSuite([ Metric.accuracy(), Metric.bleu(), Metric.bertscore(), Metric.custom("groundedness", check_groundedness) ]) test_cases = [ { "input": "What is the capital of France?", "expected": "Paris", "context": "France is a country in Europe. Paris is its capital." }, ] results = await suite.evaluate(model=your_model, test_cases=test_cases) ``` ## Automated Metrics Implementation ### BLEU Score ```python from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction def calculate_bleu(reference: str, hypothesis: str, **kwargs) -> float: """Calculate BLEU score between reference and hypothesis.""" smoothie = SmoothingFunction().method4 return sentence_bleu( [reference.split()], hypothesis.split(), smoothing_function=smoothie ) ``` ### ROUGE Score ```python from rouge_score import rouge_scorer def calculate_rouge(reference: str, hypothesis: str, **kwargs) -> dict: """Calculate ROUGE scores.""" scorer = rouge_scorer.RougeScorer( ['rouge1', 'rouge2', 'rougeL'], use_stemmer=True ) scores = scorer.score(reference, hypothesis) return { 'rouge1': scores['rouge1'].fmeasure, 'rouge2': scores['rouge2'].fmeasure, 'rougeL': scores['rougeL'].fmeasure } ``` ### BERTScore ```python from bert_score import score def calculate_bertscore( references: list[str], hypotheses: list[str], **kwargs ) -> dict: """Calculate BERTScore using pre-trained model.""" P, R, F1 = score( hypotheses, references, lang='en', model_type='microsoft/deberta-xlarge-mnli' ) return { 'precision': P.mean().item(), 'recall': R.mean().item(), 'f1': F1.mean().item() } ``` ### Custom Metrics ```python def calculate_groundedness(response: str, context: str, **kwargs) -> float: """Check if response is grounded in provided context.""" from transformers import pipeline nli = pipeline( "text-classification", model="microsoft/deberta-large-mnli" ) result = nli(f"{context} [SEP] {response}")[0] # Return confidence that response is entailed by context return result['score'] if result['label'] == 'ENTAILMENT' else 0.0 def calculate_toxicity(text: str, **kwargs) -> float: """Measure toxicity in generated text.""" from detoxify import Detoxify results = Detoxify('original').predict(text) return max(results.values()) # Return highest toxicity score def calculate_factuality(claim: str, sources: list[str], **kwargs) -> float: """Verify factual claims against sources.""" from transformers import pipeline nli = pipeline("text-classification", model="facebook/bart-large-mnli") scores = [] for source in sources: result = nli(f"{source}</s></s>{claim}")[0] if result['label'] == 'entailment': scores.append(result['score']) return max(scores) if scores else 0.0 ``` ## LLM-as-Judge Patterns ### Single Output Evaluation ```python from anthropic import Anthropic from pydantic import BaseModel, Field import json class QualityRating(BaseModel): accuracy: int = Field(ge=1, le=10, description="Factual correctness") helpfulness: int = Field(ge=1, le=10, description="Answers the question") clarity: int = Field(ge=1, le=10, description="Well-written and understandable") reasoning: str = Field(description="Brief explanation") async def llm_judge_quality( response: str, question: str, context: str = None ) -> QualityRating: """Use Claude to judge response quality.""" client = Anthropic() system = """You are an expert evaluator of AI responses. Rate responses on accuracy, helpfulness, and clarity (1-10 scale). Provide brief reasoning for your ratings.""" prompt = f"""Rate the following response: Question: {question} {f'Context: {context}' if context else ''} Response: {response} Provide ratings in JSON format: {{ "accuracy": <1-10>, "helpfulness": <1-10>, "clarity": <1-10>, "reasoning": "<brief explanation>" }}""" message = client.messages.create( model="claude-sonnet-4-6", max_tokens=500, system=system, messages=[{"role": "user", "content": prompt}] ) return QualityRating(**json.loads(message.content[0].text)) ``` ### Pairwise Comparison ```python from pydantic import BaseModel, Field from typing import Literal class ComparisonResult(BaseModel): winner: Literal["A", "B", "tie"] reasoning: str confidence: int = Field(ge=1, le=10) async def compare_responses( question: str, response_a: str, response_b: str ) -> ComparisonResult: """Compare two responses using LLM judge.""" client = Anthropic() prompt = f"""Compare these two responses and determine which is better. Question: {question} Response A: {response_a} Response B: {response_b} Consider accuracy, helpfulness, and clarity. Answer with JSON: {{ "winner": "A" or "B" or "tie", "reasoning": "<explanation>", "confidence": <1-10> }}""" message = client.messages.create( model="claude-sonnet-4-6", max_tokens=500, messages=[{"role": "user", "content": prompt}] ) return ComparisonResult(**json.loads(message.content[0].text)) ``` ### Reference-Based Evaluation ```python class ReferenceEvaluation(BaseModel): semantic_similarity: float = Field(ge=0, le=1) factual_accuracy: float = Field(ge=0, le=1) completeness: float = Field(ge=0, le=1) issues: list[str] async def evaluate_against_reference( response: str, reference: str, question: str ) -> ReferenceEvaluation: """Evaluate response against gold standard reference.""" client = Anthropic() prompt = f"""Compare the response to the reference answer. Question: {question} Reference Answer: {reference} Response to Evaluate: {response} Evaluate: 1. Semantic similarity (0-1): How similar is the meaning? 2. Factual accuracy (0-1): Are all facts correct? 3. Completeness (0-1): Does it cover all key points? 4. List any specific issues or errors. Respond in JSON: {{ "semantic_similarity": <0-1>, "factual_accuracy": <0-1>, "completeness": <0-1>, "issues": ["issue1", "issue2"] }}""" message = client.messages.create( model="claude-sonnet-4-6", max_tokens=500, messages=[{"role": "user", "content": prompt}] ) return ReferenceEvaluation(**json.loads(message.content[0].text)) ``` ## Human Evaluation Frameworks ### Annotation Guidelines ```python from dataclasses import dataclass, field from typing import Optional @dataclass class AnnotationTask: """Structure for human annotation task.""" response: str question: str context: Optional[str] = None def get_annotation_form(self) -> dict: return { "question": self.question, "context": self.context, "response": self.response, "ratings": { "accuracy": { "scale": "1-5", "description": "Is the response factually correct?" }, "relevance": { "scale": "1-5", "description": "Does it answer the question?" }, "coherence": { "scale": "1-5", "description": "Is it logically consistent?" } }, "issues": { "factual_error": False, "hallucination": False, "off_topic": False, "unsafe_content": False }, "feedback": "" } ``` ### Inter-Rater Agreement ```python from sklearn.metrics import cohen_kappa_score def calculate_agreement( rater1_scores: list[int], rater2_scores: list[int] ) -> dict: """Calculate inter-rater agreement.""" kappa = cohen_kappa_score(rater1_scores, rater2_scores) if kappa < 0: interpretation = "Poor" elif kappa < 0.2: interpretation = "Slight" elif kappa < 0.4: interpretation = "Fair" elif kappa < 0.6: interpretation = "Moderate" elif kappa < 0.8: interpretation = "Substantial" else: interpretation = "Almost Perfect" return { "kappa": kappa, "interpretation": interpretation } ``` ## A/B Testing ### Statistical Testing Framework ```python from scipy import stats import numpy as np from dataclasses import dataclass, field @dataclass class ABTest: variant_a_name: str = "A" variant_b_name: str = "B" variant_a_scores: list[float] = field(default_factory=list) variant_b_scores: list[float] = field(default_factory=list) def add_result(self, variant: str, score: float): """Add evaluation result for a variant.""" if variant == "A": self.variant_a_scores.append(score) else: self.variant_b_scores.append(score) def analyze(self, alpha: float = 0.05) -> dict: """Perform statistical analysis.""" a_scores = np.array(self.variant_a_scores) b_scores = np.array(self.variant_b_scores) # T-test t_stat, p_value = stats.ttest_ind(a_scores, b_scores) # Effect size (Cohen's d) pooled_std = np.sqrt((np.std(a_scores)**2 + np.std(b_scores)**2) / 2) cohens_d = (np.mean(b_scores) - np.mean(a_scores)) / pooled_std return { "variant_a_mean": np.mean(a_scores), "variant_b_mean": np.mean(b_scores), "difference": np.mean(b_scores) - np.mean(a_scores), "relative_improvement": (np.mean(b_scores) - np.mean(a_scores)) / np.mean(a_scores), "p_value": p_value, "statistically_significant": p_value < alpha, "cohens_d": cohens_d, "effect_size": self._interpret_cohens_d(cohens_d), "winner": self.variant_b_name if np.mean(b_scores) > np.mean(a_scores) else self.variant_a_name } @staticmethod def _interpret_cohens_d(d: float) -> str: """Interpret Cohen's d effect size.""" abs_d = abs(d) if abs_d < 0.2: return "negligible" elif abs_d < 0.5: return "small" elif abs_d < 0.8: return "medium" else: return "large" ``` ## Regression Testing ### Regression Detection ```python from dataclasses import dataclass @dataclass class RegressionResult: metric: str baseline: float current: float change: float is_regression: bool class RegressionDetector: def __init__(self, baseline_results: dict, threshold: float = 0.05): self.baseline = baseline_results self.threshold = threshold def check_for_regression(self, new_results: dict) -> dict: """Detect if new results show regression.""" regressions = [] for metric in self.baseline.keys(): baseline_score = self.baseline[metric] new_score = new_results.get(metric) if new_score is None: continue # Calculate relative change relative_change = (new_score - baseline_score) / baseline_score # Flag if significant decrease is_regression = relative_change < -self.threshold if is_regression: regressions.append(RegressionResult( metric=metric, baseline=baseline_score, current=new_score, change=relative_change, is_regression=True )) return { "has_regression": len(regressions) > 0, "regressions": regressions, "summary": f"{len(regressions)} metric(s) regressed" } ``` ## LangSmith Evaluation Integration ```python from langsmith import Client from langsmith.evaluation import evaluate, LangChainStringEvaluator # Initialize LangSmith client client = Client() # Create dataset dataset = client.create_dataset("qa_test_cases") client.create_examples( inputs=[{"question": q} for q in questions], outputs=[{"answer": a} for a in expected_answers], dataset_id=dataset.id ) # Define evaluators evaluators = [ LangChainStringEvaluator("qa"), # QA correctness LangChainStringEvaluator("context_qa"), # Context-grounded QA LangChainStringEvaluator("cot_qa"), # Chain-of-thought QA ] # Run evaluation async def target_function(inputs: dict) -> dict: result = await your_chain.ainvoke(inputs) return {"answer": result} experiment_results = await evaluate( target_function, data=dataset.name, evaluators=evaluators, experiment_prefix="v1.0.0", metadata={"model": "claude-sonnet-4-6", "version": "1.0.0"} ) print(f"Mean score: {experiment_results.aggregate_metrics['qa']['mean']}") ``` ## Benchmarking ### Running Benchmarks ```python from dataclasses import dataclass import numpy as np @dataclass class BenchmarkResult: metric: str mean: float std: float min: float max: float class BenchmarkRunner: def __init__(self, benchmark_dataset: list[dict]): self.dataset = benchmark_dataset async def run_benchmark( self, model, metrics: list[Metric] ) -> dict[str, BenchmarkResult]: """Run model on benchmark and calculate metrics.""" results = {metric.name: [] for metric in metrics} for example in self.dataset: # Generate prediction prediction = await model.predict(example["input"]) # Calculate each metric for metric in metrics: score = metric.fn( prediction=prediction, reference=example["reference"], context=example.get("context") ) results[metric.name].append(score) # Aggregate results return { metric: BenchmarkResult( metric=metric, mean=np.mean(scores), std=np.std(scores), min=min(scores), max=max(scores) ) for metric, scores in results.items() } ``` ## Resources - [LangSmith Evaluation Guide](https://docs.smith.langchain.com/evaluation) - [RAGAS Framework](https://docs.ragas.io/) - [DeepEval Library](https://docs.deepeval.com/) - [Arize Phoenix](https://docs.arize.com/phoenix/) - [HELM Benchmark](https://crfm.stanford.edu/helm/) ## Best Practices 1. **Multiple Metrics**: Use diverse metrics for comprehensive view 2. **Representative Data**: Test on real-world, diverse examples 3. **Baselines**: Always compare against baseline performance 4. **Statistical Rigor**: Use proper statistical tests for comparisons 5. **Continuous Evaluation**: Integrate into CI/CD pipeline 6. **Human Validation**: Combine automated metrics with human judgment 7. **Error Analysis**: Investigate failures to understand weaknesses 8. **Version Control**: Track evaluation results over time ## Common Pitfalls - **Single Metric Obsession**: Optimizing for one metric at the expense of others - **Small Sample Size**: Drawing conclusions from too few examples - **Data Contamination**: Testing on training data - **Ignoring Variance**: Not accounting for statistical uncertainty - **Metric Mismatch**: Using metrics not aligned with business goals - **Position Bias**: In pairwise evals, randomize order - **Overfitting Prompts**: Optimizing for test set instead of real use
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

vector-index-tuning

Optimize vector index performance for latency, recall, and memory.

coding
⭐1
# Vector Index Tuning Guide to optimizing vector indexes for production performance. ## When to Use This Skill - Tuning HNSW parameters - Implementing quantization - Optimizing memory usage - Reducing search latency - Balancing recall vs speed - Scaling to billions of vectors ## Core Concepts ### 1. Index Type Selection ``` Data Size Recommended Index ──────────────────────────────────────── < 10K vectors → Flat (exact search) 10K - 1M → HNSW 1M - 100M → HNSW + Quantization > 100M → IVF + PQ or DiskANN ``` ### 2. HNSW Parameters | Parameter | Default | Effect | | ------------------ | ------- | ---------------------------------------------------- | | **M** | 16 | Connections per node, ↑ = better recall, more memory | | **efConstruction** | 100 | Build quality, ↑ = better index, slower build | | **efSearch** | 50 | Search quality, ↑ = better recall, slower search | ### 3. Quantization Types ``` Full Precision (FP32): 4 bytes Ɨ dimensions Half Precision (FP16): 2 bytes Ɨ dimensions INT8 Scalar: 1 byte Ɨ dimensions Product Quantization: ~32-64 bytes total Binary: dimensions/8 bytes ``` ## Templates ### Template 1: HNSW Parameter Tuning ```python import numpy as np from typing import List, Tuple import time def benchmark_hnsw_parameters( vectors: np.ndarray, queries: np.ndarray, ground_truth: np.ndarray, m_values: List[int] = [8, 16, 32, 64], ef_construction_values: List[int] = [64, 128, 256], ef_search_values: List[int] = [32, 64, 128, 256] ) -> List[dict]: """Benchmark different HNSW configurations.""" import hnswlib results = [] dim = vectors.shape[1] n = vectors.shape[0] for m in m_values: for ef_construction in ef_construction_values: # Build index index = hnswlib.Index(space='cosine', dim=dim) index.init_index(max_elements=n, M=m, ef_construction=ef_construction) build_start = time.time() index.add_items(vectors) build_time = time.time() - build_start # Get memory usage memory_bytes = index.element_count * ( dim * 4 + # Vector storage m * 2 * 4 # Graph edges (approximate) ) for ef_search in ef_search_values: index.set_ef(ef_search) # Measure search search_start = time.time() labels, distances = index.knn_query(queries, k=10) search_time = time.time() - search_start # Calculate recall recall = calculate_recall(labels, ground_truth, k=10) results.append({ "M": m, "ef_construction": ef_construction, "ef_search": ef_search, "build_time_s": build_time, "search_time_ms": search_time * 1000 / len(queries), "recall@10": recall, "memory_mb": memory_bytes / 1024 / 1024 }) return results def calculate_recall(predictions: np.ndarray, ground_truth: np.ndarray, k: int) -> float: """Calculate recall@k.""" correct = 0 for pred, truth in zip(predictions, ground_truth): correct += len(set(pred[:k]) & set(truth[:k])) return correct / (len(predictions) * k) def recommend_hnsw_params( num_vectors: int, target_recall: float = 0.95, max_latency_ms: float = 10, available_memory_gb: float = 8 ) -> dict: """Recommend HNSW parameters based on requirements.""" # Base recommendations if num_vectors < 100_000: m = 16 ef_construction = 100 elif num_vectors < 1_000_000: m = 32 ef_construction = 200 else: m = 48 ef_construction = 256 # Adjust ef_search based on recall target if target_recall >= 0.99: ef_search = 256 elif target_recall >= 0.95: ef_search = 128 else: ef_search = 64 return { "M": m, "ef_construction": ef_construction, "ef_search": ef_search, "notes": f"Estimated for {num_vectors:,} vectors, {target_recall:.0%} recall" } ``` ### Template 2: Quantization Strategies ```python import numpy as np from typing import Optional class VectorQuantizer: """Quantization strategies for vector compression.""" @staticmethod def scalar_quantize_int8( vectors: np.ndarray, min_val: Optional[float] = None, max_val: Optional[float] = None ) -> Tuple[np.ndarray, dict]: """Scalar quantization to INT8.""" if min_val is None: min_val = vectors.min() if max_val is None: max_val = vectors.max() # Scale to 0-255 range scale = 255.0 / (max_val - min_val) quantized = np.clip( np.round((vectors - min_val) * scale), 0, 255 ).astype(np.uint8) params = {"min_val": min_val, "max_val": max_val, "scale": scale} return quantized, params @staticmethod def dequantize_int8( quantized: np.ndarray, params: dict ) -> np.ndarray: """Dequantize INT8 vectors.""" return quantized.astype(np.float32) / params["scale"] + params["min_val"] @staticmethod def product_quantize( vectors: np.ndarray, n_subvectors: int = 8, n_centroids: int = 256 ) -> Tuple[np.ndarray, dict]: """Product quantization for aggressive compression.""" from sklearn.cluster import KMeans n, dim = vectors.shape assert dim % n_subvectors == 0 subvector_dim = dim // n_subvectors codebooks = [] codes = np.zeros((n, n_subvectors), dtype=np.uint8) for i in range(n_subvectors): start = i * subvector_dim end = (i + 1) * subvector_dim subvectors = vectors[:, start:end] kmeans = KMeans(n_clusters=n_centroids, random_state=42) codes[:, i] = kmeans.fit_predict(subvectors) codebooks.append(kmeans.cluster_centers_) params = { "codebooks": codebooks, "n_subvectors": n_subvectors, "subvector_dim": subvector_dim } return codes, params @staticmethod def binary_quantize(vectors: np.ndarray) -> np.ndarray: """Binary quantization (sign of each dimension).""" # Convert to binary: positive = 1, negative = 0 binary = (vectors > 0).astype(np.uint8) # Pack bits into bytes n, dim = vectors.shape packed_dim = (dim + 7) // 8 packed = np.zeros((n, packed_dim), dtype=np.uint8) for i in range(dim): byte_idx = i // 8 bit_idx = i % 8 packed[:, byte_idx] |= (binary[:, i] << bit_idx) return packed def estimate_memory_usage( num_vectors: int, dimensions: int, quantization: str = "fp32", index_type: str = "hnsw", hnsw_m: int = 16 ) -> dict: """Estimate memory usage for different configurations.""" # Vector storage bytes_per_dimension = { "fp32": 4, "fp16": 2, "int8": 1, "pq": 0.05, # Approximate "binary": 0.125 } vector_bytes = num_vectors * dimensions * bytes_per_dimension[quantization] # Index overhead if index_type == "hnsw": # Each node has ~M*2 edges, each edge is 4 bytes (int32) index_bytes = num_vectors * hnsw_m * 2 * 4 elif index_type == "ivf": # Inverted lists + centroids index_bytes = num_vectors * 8 + 65536 * dimensions * 4 else: index_bytes = 0 total_bytes = vector_bytes + index_bytes return { "vector_storage_mb": vector_bytes / 1024 / 1024, "index_overhead_mb": index_bytes / 1024 / 1024, "total_mb": total_bytes / 1024 / 1024, "total_gb": total_bytes / 1024 / 1024 / 1024 } ``` ### Template 3: Qdrant Index Configuration ```python from qdrant_client import QdrantClient from qdrant_client.http import models def create_optimized_collection( client: QdrantClient, collection_name: str, vector_size: int, num_vectors: int, optimize_for: str = "balanced" # "recall", "speed", "memory" ) -> None: """Create collection with optimized settings.""" # HNSW configuration based on optimization target hnsw_configs = { "recall": models.HnswConfigDiff(m=32, ef_construct=256), "speed": models.HnswConfigDiff(m=16, ef_construct=64), "balanced": models.HnswConfigDiff(m=16, ef_construct=128), "memory": models.HnswConfigDiff(m=8, ef_construct=64) } # Quantization configuration quantization_configs = { "recall": None, # No quantization for max recall "speed": models.ScalarQuantization( scalar=models.ScalarQuantizationConfig( type=models.ScalarType.INT8, quantile=0.99, always_ram=True ) ), "balanced": models.ScalarQuantization( scalar=models.ScalarQuantizationConfig( type=models.ScalarType.INT8, quantile=0.99, always_ram=False ) ), "memory": models.ProductQuantization( product=models.ProductQuantizationConfig( compression=models.CompressionRatio.X16, always_ram=False ) ) } # Optimizer configuration optimizer_configs = { "recall": models.OptimizersConfigDiff( indexing_threshold=10000, memmap_threshold=50000 ), "speed": models.OptimizersConfigDiff( indexing_threshold=5000, memmap_threshold=20000 ), "balanced": models.OptimizersConfigDiff( indexing_threshold=20000, memmap_threshold=50000 ), "memory": models.OptimizersConfigDiff( indexing_threshold=50000, memmap_threshold=10000 # Use disk sooner ) } client.create_collection( collection_name=collection_name, vectors_config=models.VectorParams( size=vector_size, distance=models.Distance.COSINE ), hnsw_config=hnsw_configs[optimize_for], quantization_config=quantization_configs[optimize_for], optimizers_config=optimizer_configs[optimize_for] ) def tune_search_parameters( client: QdrantClient, collection_name: str, target_recall: float = 0.95 ) -> dict: """Tune search parameters for target recall.""" # Search parameter recommendations if target_recall >= 0.99: search_params = models.SearchParams( hnsw_ef=256, exact=False, quantization=models.QuantizationSearchParams( ignore=True, # Don't use quantization for search rescore=True ) ) elif target_recall >= 0.95: search_params = models.SearchParams( hnsw_ef=128, exact=False, quantization=models.QuantizationSearchParams( ignore=False, rescore=True, oversampling=2.0 ) ) else: search_params = models.SearchParams( hnsw_ef=64, exact=False, quantization=models.QuantizationSearchParams( ignore=False, rescore=False ) ) return search_params ``` ### Template 4: Performance Monitoring ```python import time from dataclasses import dataclass from typing import List import numpy as np @dataclass class SearchMetrics: latency_p50_ms: float latency_p95_ms: float latency_p99_ms: float recall: float qps: float class VectorSearchMonitor: """Monitor vector search performance.""" def __init__(self, ground_truth_fn=None): self.latencies = [] self.recalls = [] self.ground_truth_fn = ground_truth_fn def measure_search( self, search_fn, query_vectors: np.ndarray, k: int = 10, num_iterations: int = 100 ) -> SearchMetrics: """Benchmark search performance.""" latencies = [] for _ in range(num_iterations): for query in query_vectors: start = time.perf_counter() results = search_fn(query, k=k) latency = (time.perf_counter() - start) * 1000 latencies.append(latency) latencies = np.array(latencies) total_queries = num_iterations * len(query_vectors) total_time = sum(latencies) / 1000 # seconds return SearchMetrics( latency_p50_ms=np.percentile(latencies, 50), latency_p95_ms=np.percentile(latencies, 95), latency_p99_ms=np.percentile(latencies, 99), recall=self._calculate_recall(search_fn, query_vectors, k) if self.ground_truth_fn else 0, qps=total_queries / total_time ) def _calculate_recall(self, search_fn, queries: np.ndarray, k: int) -> float: """Calculate recall against ground truth.""" if not self.ground_truth_fn: return 0 correct = 0 total = 0 for query in queries: predicted = set(search_fn(query, k=k)) actual = set(self.ground_truth_fn(query, k=k)) correct += len(predicted & actual) total += k return correct / total def profile_index_build( build_fn, vectors: np.ndarray, batch_sizes: List[int] = [1000, 10000, 50000] ) -> dict: """Profile index build performance.""" results = {} for batch_size in batch_sizes: times = [] for i in range(0, len(vectors), batch_size): batch = vectors[i:i + batch_size] start = time.perf_counter() build_fn(batch) times.append(time.perf_counter() - start) results[batch_size] = { "avg_batch_time_s": np.mean(times), "vectors_per_second": batch_size / np.mean(times) } return results ``` ## Best Practices ### Do's - **Benchmark with real queries** - Synthetic may not represent production - **Monitor recall continuously** - Can degrade with data drift - **Start with defaults** - Tune only when needed - **Use quantization** - Significant memory savings - **Consider tiered storage** - Hot/cold data separation ### Don'ts - **Don't over-optimize early** - Profile first - **Don't ignore build time** - Index updates have cost - **Don't forget reindexing** - Plan for maintenance - **Don't skip warming** - Cold indexes are slow ## Resources - [HNSW Paper](https://arxiv.org/abs/1603.09320) - [Faiss Wiki](https://github.com/facebookresearch/faiss/wiki) - [ANN Benchmarks](https://ann-benchmarks.com/)
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

slo-implementation

Define and implement Service Level Indicators (SLIs) and Service

coding
⭐1
# SLO Implementation Framework for defining and implementing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. ## Purpose Implement measurable reliability targets using SLIs, SLOs, and error budgets to balance reliability with innovation velocity. ## When to Use - Define service reliability targets - Measure user-perceived reliability - Implement error budgets - Create SLO-based alerts - Track reliability goals ## SLI/SLO/SLA Hierarchy ``` SLA (Service Level Agreement) ↓ Contract with customers SLO (Service Level Objective) ↓ Internal reliability target SLI (Service Level Indicator) ↓ Actual measurement ``` ## Defining SLIs ### Common SLI Types #### 1. Availability SLI ```promql # Successful requests / Total requests sum(rate(http_requests_total{status!~"5.."}[28d])) / sum(rate(http_requests_total[28d])) ``` #### 2. Latency SLI ```promql # Requests below latency threshold / Total requests sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d])) / sum(rate(http_request_duration_seconds_count[28d])) ``` #### 3. Durability SLI ``` # Successful writes / Total writes sum(storage_writes_successful_total) / sum(storage_writes_total) ``` **Reference:** See `references/slo-definitions.md` ## Setting SLO Targets ### Availability SLO Examples | SLO % | Downtime/Month | Downtime/Year | | ------ | -------------- | ------------- | | 99% | 7.2 hours | 3.65 days | | 99.9% | 43.2 minutes | 8.76 hours | | 99.95% | 21.6 minutes | 4.38 hours | | 99.99% | 4.32 minutes | 52.56 minutes | ### Choose Appropriate SLOs **Consider:** - User expectations - Business requirements - Current performance - Cost of reliability - Competitor benchmarks **Example SLOs:** ```yaml slos: - name: api_availability target: 99.9 window: 28d sli: | sum(rate(http_requests_total{status!~"5.."}[28d])) / sum(rate(http_requests_total[28d])) - name: api_latency_p95 target: 99 window: 28d sli: | sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d])) / sum(rate(http_request_duration_seconds_count[28d])) ``` ## Error Budget Calculation ### Error Budget Formula ``` Error Budget = 1 - SLO Target ``` **Example:** - SLO: 99.9% availability - Error Budget: 0.1% = 43.2 minutes/month - Current Error: 0.05% = 21.6 minutes/month - Remaining Budget: 50% ### Error Budget Policy ```yaml error_budget_policy: - remaining_budget: 100% action: Normal development velocity - remaining_budget: 50% action: Consider postponing risky changes - remaining_budget: 10% action: Freeze non-critical changes - remaining_budget: 0% action: Feature freeze, focus on reliability ``` **Reference:** See `references/error-budget.md` ## SLO Implementation ### Prometheus Recording Rules ```yaml # SLI Recording Rules groups: - name: sli_rules interval: 30s rules: # Availability SLI - record: sli:http_availability:ratio expr: | sum(rate(http_requests_total{status!~"5.."}[28d])) / sum(rate(http_requests_total[28d])) # Latency SLI (requests < 500ms) - record: sli:http_latency:ratio expr: | sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d])) / sum(rate(http_request_duration_seconds_count[28d])) - name: slo_rules interval: 5m rules: # SLO compliance (1 = meeting SLO, 0 = violating) - record: slo:http_availability:compliance expr: sli:http_availability:ratio >= bool 0.999 - record: slo:http_latency:compliance expr: sli:http_latency:ratio >= bool 0.99 # Error budget remaining (percentage) - record: slo:http_availability:error_budget_remaining expr: | (sli:http_availability:ratio - 0.999) / (1 - 0.999) * 100 # Error budget burn rate - record: slo:http_availability:burn_rate_5m expr: | (1 - ( sum(rate(http_requests_total{status!~"5.."}[5m])) / sum(rate(http_requests_total[5m])) )) / (1 - 0.999) ``` ### SLO Alerting Rules ```yaml groups: - name: slo_alerts interval: 1m rules: # Fast burn: 14.4x rate, 1 hour window # Consumes 2% error budget in 1 hour - alert: SLOErrorBudgetBurnFast expr: | slo:http_availability:burn_rate_1h > 14.4 and slo:http_availability:burn_rate_5m > 14.4 for: 2m labels: severity: critical annotations: summary: "Fast error budget burn detected" description: "Error budget burning at {{ $value }}x rate" # Slow burn: 6x rate, 6 hour window # Consumes 5% error budget in 6 hours - alert: SLOErrorBudgetBurnSlow expr: | slo:http_availability:burn_rate_6h > 6 and slo:http_availability:burn_rate_30m > 6 for: 15m labels: severity: warning annotations: summary: "Slow error budget burn detected" description: "Error budget burning at {{ $value }}x rate" # Error budget exhausted - alert: SLOErrorBudgetExhausted expr: slo:http_availability:error_budget_remaining < 0 for: 5m labels: severity: critical annotations: summary: "SLO error budget exhausted" description: "Error budget remaining: {{ $value }}%" ``` ## SLO Dashboard **Grafana Dashboard Structure:** ``` ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā” │ SLO Compliance (Current) │ │ āœ“ 99.95% (Target: 99.9%) │ ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤ │ Error Budget Remaining: 65% │ │ ā–ˆā–ˆā–ˆā–ˆā–ˆā–ˆā–ˆā–ˆā–‘ā–‘ 65% │ ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤ │ SLI Trend (28 days) │ │ [Time series graph] │ ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤ │ Burn Rate Analysis │ │ [Burn rate by time window] │ ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜ ``` **Example Queries:** ```promql # Current SLO compliance sli:http_availability:ratio * 100 # Error budget remaining slo:http_availability:error_budget_remaining # Days until error budget exhausted (at current burn rate) (slo:http_availability:error_budget_remaining / 100) * 28 / (1 - sli:http_availability:ratio) * (1 - 0.999) ``` ## Multi-Window Burn Rate Alerts ```yaml # Combination of short and long windows reduces false positives rules: - alert: SLOBurnRateHigh expr: | ( slo:http_availability:burn_rate_1h > 14.4 and slo:http_availability:burn_rate_5m > 14.4 ) or ( slo:http_availability:burn_rate_6h > 6 and slo:http_availability:burn_rate_30m > 6 ) labels: severity: critical ``` ## SLO Review Process ### Weekly Review - Current SLO compliance - Error budget status - Trend analysis - Incident impact ### Monthly Review - SLO achievement - Error budget usage - Incident postmortems - SLO adjustments ### Quarterly Review - SLO relevance - Target adjustments - Process improvements - Tooling enhancements ## Best Practices 1. **Start with user-facing services** 2. **Use multiple SLIs** (availability, latency, etc.) 3. **Set achievable SLOs** (don't aim for 100%) 4. **Implement multi-window alerts** to reduce noise 5. **Track error budget** consistently 6. **Review SLOs regularly** 7. **Document SLO decisions** 8. **Align with business goals** 9. **Automate SLO reporting** 10. **Use SLOs for prioritization** ## Reference Files - `assets/slo-template.md` - SLO definition template - `references/slo-definitions.md` - SLO definition patterns - `references/error-budget.md` - Error budget calculations ## Related Skills - `prometheus-configuration` - For metric collection - `grafana-dashboards` - For SLO visualization
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

data-storytelling

Transform data into compelling narratives using visualization,

data
⭐1
# Data Storytelling Transform raw data into compelling narratives that drive decisions and inspire action. ## When to Use This Skill - Presenting analytics to executives - Creating quarterly business reviews - Building investor presentations - Writing data-driven reports - Communicating insights to non-technical audiences - Making recommendations based on data ## Core Concepts ### 1. Story Structure ``` Setup → Conflict → Resolution Setup: Context and baseline Conflict: The problem or opportunity Resolution: Insights and recommendations ``` ### 2. Narrative Arc ``` 1. Hook: Grab attention with surprising insight 2. Context: Establish the baseline 3. Rising Action: Build through data points 4. Climax: The key insight 5. Resolution: Recommendations 6. Call to Action: Next steps ``` ### 3. Three Pillars | Pillar | Purpose | Components | | ------------- | -------- | -------------------------------- | | **Data** | Evidence | Numbers, trends, comparisons | | **Narrative** | Meaning | Context, causation, implications | | **Visuals** | Clarity | Charts, diagrams, highlights | ## Story Frameworks ### Framework 1: The Problem-Solution Story ```markdown # Customer Churn Analysis ## The Hook "We're losing $2.4M annually to preventable churn." ## The Context - Current churn rate: 8.5% (industry average: 5%) - Average customer lifetime value: $4,800 - 500 customers churned last quarter ## The Problem Analysis of churned customers reveals a pattern: - 73% churned within first 90 days - Common factor: < 3 support interactions - Low feature adoption in first month ## The Insight [Show engagement curve visualization] Customers who don't engage in the first 14 days are 4x more likely to churn. ## The Solution 1. Implement 14-day onboarding sequence 2. Proactive outreach at day 7 3. Feature adoption tracking ## Expected Impact - Reduce early churn by 40% - Save $960K annually - Payback period: 3 months ## Call to Action Approve $50K budget for onboarding automation. ``` ### Framework 2: The Trend Story ```markdown # Q4 Performance Analysis ## Where We Started Q3 ended with $1.2M MRR, 15% below target. Team morale was low after missed goals. ## What Changed [Timeline visualization] - Oct: Launched self-serve pricing - Nov: Reduced friction in signup - Dec: Added customer success calls ## The Transformation [Before/after comparison chart] | Metric | Q3 | Q4 | Change | |----------------|--------|--------|--------| | Trial → Paid | 8% | 15% | +87% | | Time to Value | 14 days| 5 days | -64% | | Expansion Rate | 2% | 8% | +300% | ## Key Insight Self-serve + high-touch creates compound growth. Customers who self-serve AND get a success call have 3x higher expansion rate. ## Going Forward Double down on hybrid model. Target: $1.8M MRR by Q2. ``` ### Framework 3: The Comparison Story ```markdown # Market Opportunity Analysis ## The Question Should we expand into EMEA or APAC first? ## The Comparison [Side-by-side market analysis] ### EMEA - Market size: $4.2B - Growth rate: 8% - Competition: High - Regulatory: Complex (GDPR) - Language: Multiple ### APAC - Market size: $3.8B - Growth rate: 15% - Competition: Moderate - Regulatory: Varied - Language: Multiple ## The Analysis [Weighted scoring matrix visualization] | Factor | Weight | EMEA Score | APAC Score | | ----------- | ------ | ---------- | ---------- | | Market Size | 25% | 5 | 4 | | Growth | 30% | 3 | 5 | | Competition | 20% | 2 | 4 | | Ease | 25% | 2 | 3 | | **Total** | | **2.9** | **4.1** | ## The Recommendation APAC first. Higher growth, less competition. Start with Singapore hub (English, business-friendly). Enter EMEA in Year 2 with localization ready. ## Risk Mitigation - Timezone coverage: Hire 24/7 support - Cultural fit: Local partnerships - Payment: Multi-currency from day 1 ``` ## Visualization Techniques ### Technique 1: Progressive Reveal ```markdown Start simple, add layers: Slide 1: "Revenue is growing" [single line chart] Slide 2: "But growth is slowing" [add growth rate overlay] Slide 3: "Driven by one segment" [add segment breakdown] Slide 4: "Which is saturating" [add market share] Slide 5: "We need new segments" [add opportunity zones] ``` ### Technique 2: Contrast and Compare ```markdown Before/After: ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā” │ BEFORE │ AFTER │ │ │ │ │ Process: 5 days│ Process: 1 day │ │ Errors: 15% │ Errors: 2% │ │ Cost: $50/unit │ Cost: $20/unit │ ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”“ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜ This/That (emphasize difference): ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā” │ CUSTOMER A vs B │ │ ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā” ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā” │ │ │ ā–ˆā–ˆā–ˆā–ˆā–ˆā–ˆā–ˆā–ˆ │ │ ā–ˆā–ˆ │ │ │ │ $45,000 │ │ $8,000 │ │ │ │ LTV │ │ LTV │ │ │ ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜ ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜ │ │ Onboarded No onboarding │ ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜ ``` ### Technique 3: Annotation and Highlight ```python import matplotlib.pyplot as plt import pandas as pd fig, ax = plt.subplots(figsize=(12, 6)) # Plot the main data ax.plot(dates, revenue, linewidth=2, color='#2E86AB') # Add annotation for key events ax.annotate( 'Product Launch\n+32% spike', xy=(launch_date, launch_revenue), xytext=(launch_date, launch_revenue * 1.2), fontsize=10, arrowprops=dict(arrowstyle='->', color='#E63946'), color='#E63946' ) # Highlight a region ax.axvspan(growth_start, growth_end, alpha=0.2, color='green', label='Growth Period') # Add threshold line ax.axhline(y=target, color='gray', linestyle='--', label=f'Target: ${target:,.0f}') ax.set_title('Revenue Growth Story', fontsize=14, fontweight='bold') ax.legend() ``` ## Presentation Templates ### Template 1: Executive Summary Slide ``` ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā” │ KEY INSIGHT │ │ ══════════════════════════════════════════════════════════│ │ │ │ "Customers who complete onboarding in week 1 │ │ have 3x higher lifetime value" │ │ │ ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤ │ │ │ │ THE DATA │ THE IMPLICATION │ │ │ │ │ Week 1 completers: │ āœ“ Prioritize onboarding UX │ │ • LTV: $4,500 │ āœ“ Add day-1 success milestones │ │ • Retention: 85% │ āœ“ Proactive week-1 outreach │ │ • NPS: 72 │ │ │ │ Investment: $75K │ │ Others: │ Expected ROI: 8x │ │ • LTV: $1,500 │ │ │ • Retention: 45% │ │ │ • NPS: 34 │ │ │ │ │ ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”“ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜ ``` ### Template 2: Data Story Flow ``` Slide 1: THE HEADLINE "We can grow 40% faster by fixing onboarding" Slide 2: THE CONTEXT Current state metrics Industry benchmarks Gap analysis Slide 3: THE DISCOVERY What the data revealed Surprising finding Pattern identification Slide 4: THE DEEP DIVE Root cause analysis Segment breakdowns Statistical significance Slide 5: THE RECOMMENDATION Proposed actions Resource requirements Timeline Slide 6: THE IMPACT Expected outcomes ROI calculation Risk assessment Slide 7: THE ASK Specific request Decision needed Next steps ``` ### Template 3: One-Page Dashboard Story ```markdown # Monthly Business Review: January 2024 ## THE HEADLINE Revenue up 15% but CAC increasing faster than LTV ## KEY METRICS AT A GLANCE ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā” │ MRR │ NRR │ CAC │ LTV │ │ $125K │ 108% │ $450 │ $2,200 │ │ ā–²15% │ ā–²3% │ ā–²22% │ ā–²8% │ ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”“ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”“ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”“ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜ ## WHAT'S WORKING āœ“ Enterprise segment growing 25% MoM āœ“ Referral program driving 30% of new logos āœ“ Support satisfaction at all-time high (94%) ## WHAT NEEDS ATTENTION āœ— SMB acquisition cost up 40% āœ— Trial conversion down 5 points āœ— Time-to-value increased by 3 days ## ROOT CAUSE [Mini chart showing SMB vs Enterprise CAC trend] SMB paid ads becoming less efficient. CPC up 35% while conversion flat. ## RECOMMENDATION 1. Shift $20K/mo from paid to content 2. Launch SMB self-serve trial 3. A/B test shorter onboarding ## NEXT MONTH'S FOCUS - Launch content marketing pilot - Complete self-serve MVP - Reduce time-to-value to < 7 days ``` ## Writing Techniques ### Headlines That Work ```markdown BAD: "Q4 Sales Analysis" GOOD: "Q4 Sales Beat Target by 23% - Here's Why" BAD: "Customer Churn Report" GOOD: "We're Losing $2.4M to Preventable Churn" BAD: "Marketing Performance" GOOD: "Content Marketing Delivers 4x ROI vs. Paid" Formula: [Specific Number] + [Business Impact] + [Actionable Context] ``` ### Transition Phrases ```markdown Building the narrative: • "This leads us to ask..." • "When we dig deeper..." • "The pattern becomes clear when..." • "Contrast this with..." Introducing insights: • "The data reveals..." • "What surprised us was..." • "The inflection point came when..." • "The key finding is..." Moving to action: • "This insight suggests..." • "Based on this analysis..." • "The implication is clear..." • "Our recommendation is..." ``` ### Handling Uncertainty ```markdown Acknowledge limitations: • "With 95% confidence, we can say..." • "The sample size of 500 shows..." • "While correlation is strong, causation requires..." • "This trend holds for [segment], though [caveat]..." Present ranges: • "Impact estimate: $400K-$600K" • "Confidence interval: 15-20% improvement" • "Best case: X, Conservative: Y" ``` ## Best Practices ### Do's - **Start with the "so what"** - Lead with insight - **Use the rule of three** - Three points, three comparisons - **Show, don't tell** - Let data speak - **Make it personal** - Connect to audience goals - **End with action** - Clear next steps ### Don'ts - **Don't data dump** - Curate ruthlessly - **Don't bury the insight** - Front-load key findings - **Don't use jargon** - Match audience vocabulary - **Don't show methodology first** - Context, then method - **Don't forget the narrative** - Numbers need meaning ## Resources - [Storytelling with Data (Cole Nussbaumer)](https://www.storytellingwithdata.com/) - [The Pyramid Principle (Barbara Minto)](https://www.amazon.com/Pyramid-Principle-Logic-Writing-Thinking/dp/0273710516) - [Resonate (Nancy Duarte)](https://www.duarte.com/resonate/)
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

risk-metrics-calculation

Calculate portfolio risk metrics including VaR, CVaR, Sharpe,

coding
⭐1
# Risk Metrics Calculation Comprehensive risk measurement toolkit for portfolio management, including Value at Risk, Expected Shortfall, and drawdown analysis. ## When to Use This Skill - Measuring portfolio risk - Implementing risk limits - Building risk dashboards - Calculating risk-adjusted returns - Setting position sizes - Regulatory reporting ## Core Concepts ### 1. Risk Metric Categories | Category | Metrics | Use Case | | ----------------- | --------------- | -------------------- | | **Volatility** | Std Dev, Beta | General risk | | **Tail Risk** | VaR, CVaR | Extreme losses | | **Drawdown** | Max DD, Calmar | Capital preservation | | **Risk-Adjusted** | Sharpe, Sortino | Performance | ### 2. Time Horizons ``` Intraday: Minute/hourly VaR for day traders Daily: Standard risk reporting Weekly: Rebalancing decisions Monthly: Performance attribution Annual: Strategic allocation ``` ## Implementation ### Pattern 1: Core Risk Metrics ```python import numpy as np import pandas as pd from scipy import stats from typing import Dict, Optional, Tuple class RiskMetrics: """Core risk metric calculations.""" def __init__(self, returns: pd.Series, rf_rate: float = 0.02): """ Args: returns: Series of periodic returns rf_rate: Annual risk-free rate """ self.returns = returns self.rf_rate = rf_rate self.ann_factor = 252 # Trading days per year # Volatility Metrics def volatility(self, annualized: bool = True) -> float: """Standard deviation of returns.""" vol = self.returns.std() if annualized: vol *= np.sqrt(self.ann_factor) return vol def downside_deviation(self, threshold: float = 0, annualized: bool = True) -> float: """Standard deviation of returns below threshold.""" downside = self.returns[self.returns < threshold] if len(downside) == 0: return 0.0 dd = downside.std() if annualized: dd *= np.sqrt(self.ann_factor) return dd def beta(self, market_returns: pd.Series) -> float: """Beta relative to market.""" aligned = pd.concat([self.returns, market_returns], axis=1).dropna() if len(aligned) < 2: return np.nan cov = np.cov(aligned.iloc[:, 0], aligned.iloc[:, 1]) return cov[0, 1] / cov[1, 1] if cov[1, 1] != 0 else 0 # Value at Risk def var_historical(self, confidence: float = 0.95) -> float: """Historical VaR at confidence level.""" return -np.percentile(self.returns, (1 - confidence) * 100) def var_parametric(self, confidence: float = 0.95) -> float: """Parametric VaR assuming normal distribution.""" z_score = stats.norm.ppf(confidence) return self.returns.mean() - z_score * self.returns.std() def var_cornish_fisher(self, confidence: float = 0.95) -> float: """VaR with Cornish-Fisher expansion for non-normality.""" z = stats.norm.ppf(confidence) s = stats.skew(self.returns) # Skewness k = stats.kurtosis(self.returns) # Excess kurtosis # Cornish-Fisher expansion z_cf = (z + (z**2 - 1) * s / 6 + (z**3 - 3*z) * k / 24 - (2*z**3 - 5*z) * s**2 / 36) return -(self.returns.mean() + z_cf * self.returns.std()) # Conditional VaR (Expected Shortfall) def cvar(self, confidence: float = 0.95) -> float: """Expected Shortfall / CVaR / Average VaR.""" var = self.var_historical(confidence) return -self.returns[self.returns <= -var].mean() # Drawdown Analysis def drawdowns(self) -> pd.Series: """Calculate drawdown series.""" cumulative = (1 + self.returns).cumprod() running_max = cumulative.cummax() return (cumulative - running_max) / running_max def max_drawdown(self) -> float: """Maximum drawdown.""" return self.drawdowns().min() def avg_drawdown(self) -> float: """Average drawdown.""" dd = self.drawdowns() return dd[dd < 0].mean() if (dd < 0).any() else 0 def drawdown_duration(self) -> Dict[str, int]: """Drawdown duration statistics.""" dd = self.drawdowns() in_drawdown = dd < 0 # Find drawdown periods drawdown_starts = in_drawdown & ~in_drawdown.shift(1).fillna(False) drawdown_ends = ~in_drawdown & in_drawdown.shift(1).fillna(False) durations = [] current_duration = 0 for i in range(len(dd)): if in_drawdown.iloc[i]: current_duration += 1 elif current_duration > 0: durations.append(current_duration) current_duration = 0 if current_duration > 0: durations.append(current_duration) return { "max_duration": max(durations) if durations else 0, "avg_duration": np.mean(durations) if durations else 0, "current_duration": current_duration } # Risk-Adjusted Returns def sharpe_ratio(self) -> float: """Annualized Sharpe ratio.""" excess_return = self.returns.mean() * self.ann_factor - self.rf_rate vol = self.volatility(annualized=True) return excess_return / vol if vol > 0 else 0 def sortino_ratio(self) -> float: """Sortino ratio using downside deviation.""" excess_return = self.returns.mean() * self.ann_factor - self.rf_rate dd = self.downside_deviation(threshold=0, annualized=True) return excess_return / dd if dd > 0 else 0 def calmar_ratio(self) -> float: """Calmar ratio (return / max drawdown).""" annual_return = (1 + self.returns).prod() ** (self.ann_factor / len(self.returns)) - 1 max_dd = abs(self.max_drawdown()) return annual_return / max_dd if max_dd > 0 else 0 def omega_ratio(self, threshold: float = 0) -> float: """Omega ratio.""" returns_above = self.returns[self.returns > threshold] - threshold returns_below = threshold - self.returns[self.returns <= threshold] if returns_below.sum() == 0: return np.inf return returns_above.sum() / returns_below.sum() # Information Ratio def information_ratio(self, benchmark_returns: pd.Series) -> float: """Information ratio vs benchmark.""" active_returns = self.returns - benchmark_returns tracking_error = active_returns.std() * np.sqrt(self.ann_factor) active_return = active_returns.mean() * self.ann_factor return active_return / tracking_error if tracking_error > 0 else 0 # Summary def summary(self) -> Dict[str, float]: """Generate comprehensive risk summary.""" dd_stats = self.drawdown_duration() return { # Returns "total_return": (1 + self.returns).prod() - 1, "annual_return": (1 + self.returns).prod() ** (self.ann_factor / len(self.returns)) - 1, # Volatility "annual_volatility": self.volatility(), "downside_deviation": self.downside_deviation(), # VaR & CVaR "var_95_historical": self.var_historical(0.95), "var_99_historical": self.var_historical(0.99), "cvar_95": self.cvar(0.95), # Drawdowns "max_drawdown": self.max_drawdown(), "avg_drawdown": self.avg_drawdown(), "max_drawdown_duration": dd_stats["max_duration"], # Risk-Adjusted "sharpe_ratio": self.sharpe_ratio(), "sortino_ratio": self.sortino_ratio(), "calmar_ratio": self.calmar_ratio(), "omega_ratio": self.omega_ratio(), # Distribution "skewness": stats.skew(self.returns), "kurtosis": stats.kurtosis(self.returns), } ``` ### Pattern 2: Portfolio Risk ```python class PortfolioRisk: """Portfolio-level risk calculations.""" def __init__( self, returns: pd.DataFrame, weights: Optional[pd.Series] = None ): """ Args: returns: DataFrame with asset returns (columns = assets) weights: Portfolio weights (default: equal weight) """ self.returns = returns self.weights = weights if weights is not None else \ pd.Series(1/len(returns.columns), index=returns.columns) self.ann_factor = 252 def portfolio_return(self) -> float: """Weighted portfolio return.""" return (self.returns @ self.weights).mean() * self.ann_factor def portfolio_volatility(self) -> float: """Portfolio volatility.""" cov_matrix = self.returns.cov() * self.ann_factor port_var = self.weights @ cov_matrix @ self.weights return np.sqrt(port_var) def marginal_risk_contribution(self) -> pd.Series: """Marginal contribution to risk by asset.""" cov_matrix = self.returns.cov() * self.ann_factor port_vol = self.portfolio_volatility() # Marginal contribution mrc = (cov_matrix @ self.weights) / port_vol return mrc def component_risk(self) -> pd.Series: """Component contribution to total risk.""" mrc = self.marginal_risk_contribution() return self.weights * mrc def risk_parity_weights(self, target_vol: float = None) -> pd.Series: """Calculate risk parity weights.""" from scipy.optimize import minimize n = len(self.returns.columns) cov_matrix = self.returns.cov() * self.ann_factor def risk_budget_objective(weights): port_vol = np.sqrt(weights @ cov_matrix @ weights) mrc = (cov_matrix @ weights) / port_vol rc = weights * mrc target_rc = port_vol / n # Equal risk contribution return np.sum((rc - target_rc) ** 2) constraints = [ {"type": "eq", "fun": lambda w: np.sum(w) - 1}, # Weights sum to 1 ] bounds = [(0.01, 1.0) for _ in range(n)] # Min 1%, max 100% x0 = np.array([1/n] * n) result = minimize( risk_budget_objective, x0, method="SLSQP", bounds=bounds, constraints=constraints ) return pd.Series(result.x, index=self.returns.columns) def correlation_matrix(self) -> pd.DataFrame: """Asset correlation matrix.""" return self.returns.corr() def diversification_ratio(self) -> float: """Diversification ratio (higher = more diversified).""" asset_vols = self.returns.std() * np.sqrt(self.ann_factor) weighted_vol = (self.weights * asset_vols).sum() port_vol = self.portfolio_volatility() return weighted_vol / port_vol if port_vol > 0 else 1 def tracking_error(self, benchmark_returns: pd.Series) -> float: """Tracking error vs benchmark.""" port_returns = self.returns @ self.weights active_returns = port_returns - benchmark_returns return active_returns.std() * np.sqrt(self.ann_factor) def conditional_correlation( self, threshold_percentile: float = 10 ) -> pd.DataFrame: """Correlation during stress periods.""" port_returns = self.returns @ self.weights threshold = np.percentile(port_returns, threshold_percentile) stress_mask = port_returns <= threshold return self.returns[stress_mask].corr() ``` ### Pattern 3: Rolling Risk Metrics ```python class RollingRiskMetrics: """Rolling window risk calculations.""" def __init__(self, returns: pd.Series, window: int = 63): """ Args: returns: Return series window: Rolling window size (default: 63 = ~3 months) """ self.returns = returns self.window = window def rolling_volatility(self, annualized: bool = True) -> pd.Series: """Rolling volatility.""" vol = self.returns.rolling(self.window).std() if annualized: vol *= np.sqrt(252) return vol def rolling_sharpe(self, rf_rate: float = 0.02) -> pd.Series: """Rolling Sharpe ratio.""" rolling_return = self.returns.rolling(self.window).mean() * 252 rolling_vol = self.rolling_volatility() return (rolling_return - rf_rate) / rolling_vol def rolling_var(self, confidence: float = 0.95) -> pd.Series: """Rolling historical VaR.""" return self.returns.rolling(self.window).apply( lambda x: -np.percentile(x, (1 - confidence) * 100), raw=True ) def rolling_max_drawdown(self) -> pd.Series: """Rolling maximum drawdown.""" def max_dd(returns): cumulative = (1 + returns).cumprod() running_max = cumulative.cummax() drawdowns = (cumulative - running_max) / running_max return drawdowns.min() return self.returns.rolling(self.window).apply(max_dd, raw=False) def rolling_beta(self, market_returns: pd.Series) -> pd.Series: """Rolling beta vs market.""" def calc_beta(window_data): port_ret = window_data.iloc[:, 0] mkt_ret = window_data.iloc[:, 1] cov = np.cov(port_ret, mkt_ret) return cov[0, 1] / cov[1, 1] if cov[1, 1] != 0 else 0 combined = pd.concat([self.returns, market_returns], axis=1) return combined.rolling(self.window).apply( lambda x: calc_beta(x.to_frame()), raw=False ).iloc[:, 0] def volatility_regime( self, low_threshold: float = 0.10, high_threshold: float = 0.20 ) -> pd.Series: """Classify volatility regime.""" vol = self.rolling_volatility() def classify(v): if v < low_threshold: return "low" elif v > high_threshold: return "high" else: return "normal" return vol.apply(classify) ``` ### Pattern 4: Stress Testing ```python class StressTester: """Historical and hypothetical stress testing.""" # Historical crisis periods HISTORICAL_SCENARIOS = { "2008_financial_crisis": ("2008-09-01", "2009-03-31"), "2020_covid_crash": ("2020-02-19", "2020-03-23"), "2022_rate_hikes": ("2022-01-01", "2022-10-31"), "dot_com_bust": ("2000-03-01", "2002-10-01"), "flash_crash_2010": ("2010-05-06", "2010-05-06"), } def __init__(self, returns: pd.Series, weights: pd.Series = None): self.returns = returns self.weights = weights def historical_stress_test( self, scenario_name: str, historical_data: pd.DataFrame ) -> Dict[str, float]: """Test portfolio against historical crisis period.""" if scenario_name not in self.HISTORICAL_SCENARIOS: raise ValueError(f"Unknown scenario: {scenario_name}") start, end = self.HISTORICAL_SCENARIOS[scenario_name] # Get returns during crisis crisis_returns = historical_data.loc[start:end] if self.weights is not None: port_returns = (crisis_returns @ self.weights) else: port_returns = crisis_returns total_return = (1 + port_returns).prod() - 1 max_dd = self._calculate_max_dd(port_returns) worst_day = port_returns.min() return { "scenario": scenario_name, "period": f"{start} to {end}", "total_return": total_return, "max_drawdown": max_dd, "worst_day": worst_day, "volatility": port_returns.std() * np.sqrt(252) } def hypothetical_stress_test( self, shocks: Dict[str, float] ) -> float: """ Test portfolio against hypothetical shocks. Args: shocks: Dict of {asset: shock_return} """ if self.weights is None: raise ValueError("Weights required for hypothetical stress test") total_impact = 0 for asset, shock in shocks.items(): if asset in self.weights.index: total_impact += self.weights[asset] * shock return total_impact def monte_carlo_stress( self, n_simulations: int = 10000, horizon_days: int = 21, vol_multiplier: float = 2.0 ) -> Dict[str, float]: """Monte Carlo stress test with elevated volatility.""" mean = self.returns.mean() vol = self.returns.std() * vol_multiplier simulations = np.random.normal( mean, vol, (n_simulations, horizon_days) ) total_returns = (1 + simulations).prod(axis=1) - 1 return { "expected_loss": -total_returns.mean(), "var_95": -np.percentile(total_returns, 5), "var_99": -np.percentile(total_returns, 1), "worst_case": -total_returns.min(), "prob_10pct_loss": (total_returns < -0.10).mean() } def _calculate_max_dd(self, returns: pd.Series) -> float: cumulative = (1 + returns).cumprod() running_max = cumulative.cummax() drawdowns = (cumulative - running_max) / running_max return drawdowns.min() ``` ## Quick Reference ```python # Daily usage metrics = RiskMetrics(returns) print(f"Sharpe: {metrics.sharpe_ratio():.2f}") print(f"Max DD: {metrics.max_drawdown():.2%}") print(f"VaR 95%: {metrics.var_historical(0.95):.2%}") # Full summary summary = metrics.summary() for metric, value in summary.items(): print(f"{metric}: {value:.4f}") ``` ## Best Practices ### Do's - **Use multiple metrics** - No single metric captures all risk - **Consider tail risk** - VaR isn't enough, use CVaR - **Rolling analysis** - Risk changes over time - **Stress test** - Historical and hypothetical - **Document assumptions** - Distribution, lookback, etc. ### Don'ts - **Don't rely on VaR alone** - Underestimates tail risk - **Don't assume normality** - Returns are fat-tailed - **Don't ignore correlation** - Increases in stress - **Don't use short lookbacks** - Miss regime changes - **Don't forget transaction costs** - Affects realized risk ## Resources - [Risk Management and Financial Institutions (John Hull)](https://www.amazon.com/Risk-Management-Financial-Institutions-5th/dp/1119448115) - [Quantitative Risk Management (McNeil, Frey, Embrechts)](https://www.amazon.com/Quantitative-Risk-Management-Techniques-Princeton/dp/0691166277) - [pyfolio Documentation](https://quantopian.github.io/pyfolio/)
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

startup-financial-modeling

This skill should be used when the user asks to "create financial

business
⭐1
# Startup Financial Modeling Build comprehensive 3-5 year financial models with revenue projections, cost structures, cash flow analysis, and scenario planning for early-stage startups. ## Overview Financial modeling provides the quantitative foundation for startup strategy, fundraising, and operational planning. Create realistic projections using cohort-based revenue modeling, detailed cost structures, and scenario analysis to support decision-making and investor presentations. ## Core Components ### Revenue Model **Cohort-Based Projections:** Build revenue from customer acquisition and retention by cohort. **Formula:** ``` MRR = Ī£ (Cohort Size Ɨ Retention Rate Ɨ ARPU) ARR = MRR Ɨ 12 ``` **Key Inputs:** - Monthly new customer acquisitions - Customer retention rates by month - Average revenue per user (ARPU) - Pricing and packaging assumptions - Expansion revenue (upsells, cross-sells) ### Cost Structure **Operating Expenses Categories:** 1. **Cost of Goods Sold (COGS)** - Hosting and infrastructure - Payment processing fees - Customer support (variable portion) - Third-party services per customer 2. **Sales & Marketing (S&M)** - Customer acquisition cost (CAC) - Marketing programs and advertising - Sales team compensation - Marketing tools and software 3. **Research & Development (R&D)** - Engineering team compensation - Product management - Design and UX - Development tools and infrastructure 4. **General & Administrative (G&A)** - Executive team - Finance, legal, HR - Office and facilities - Insurance and compliance ### Cash Flow Analysis **Components:** - Beginning cash balance - Cash inflows (revenue, fundraising) - Cash outflows (operating expenses, CapEx) - Ending cash balance - Monthly burn rate - Runway (months of cash remaining) **Formula:** ``` Runway = Current Cash Balance / Monthly Burn Rate Monthly Burn = Monthly Revenue - Monthly Expenses ``` ### Headcount Planning **Role-Based Hiring Plan:** Track headcount by department and role. **Key Metrics:** - Fully-loaded cost per employee - Revenue per employee - Headcount by department (% of total) **Typical Ratios (Early-Stage SaaS):** - Engineering: 40-50% - Sales & Marketing: 25-35% - G&A: 10-15% - Customer Success: 5-10% ## Financial Model Structure ### Three-Scenario Framework **Conservative Scenario (P10):** - Slower customer acquisition - Lower pricing or conversion - Higher churn rates - Extended sales cycles - Used for cash management **Base Scenario (P50):** - Most likely outcomes - Realistic assumptions - Primary planning scenario - Used for board reporting **Optimistic Scenario (P90):** - Faster growth - Better unit economics - Lower churn - Used for upside planning ### Time Horizon **Detailed Projections: 3 Years** - Monthly detail for Year 1 - Monthly detail for Year 2 - Quarterly detail for Year 3 **High-Level Projections: Years 4-5** - Annual projections - Key metrics only - Support long-term planning ## Step-by-Step Process ### Step 1: Define Business Model Clarify revenue model and pricing. **SaaS Model:** - Subscription pricing tiers - Annual vs. monthly contracts - Free trial or freemium approach - Expansion revenue strategy **Marketplace Model:** - GMV projections - Take rate (% of transactions) - Buyer and seller economics - Transaction frequency **Transactional Model:** - Transaction volume - Revenue per transaction - Frequency and seasonality ### Step 2: Build Revenue Projections Use cohort-based methodology for accuracy. **Monthly Customer Acquisition:** Define new customers acquired each month. **Retention Curve:** Model customer retention over time. **Typical SaaS Retention:** - Month 1: 100% - Month 3: 90% - Month 6: 85% - Month 12: 75% - Month 24: 70% **Revenue Calculation:** For each cohort, calculate retained customers Ɨ ARPU for each month. ### Step 3: Model Cost Structure Break down costs by category and behavior. **Fixed vs. Variable:** - Fixed: Salaries, software, rent - Variable: Hosting, payment processing, support **Scaling Assumptions:** - COGS as % of revenue - S&M as % of revenue (CAC payback) - R&D growth rate - G&A as % of total expenses ### Step 4: Create Hiring Plan Model headcount growth by role and department. **Inputs:** - Starting headcount - Hiring velocity by role - Fully-loaded compensation by role - Benefits and taxes (typically 1.3-1.4x salary) **Example:** ``` Engineer: $150K salary Ɨ 1.35 = $202K fully-loaded Sales Rep: $100K OTE Ɨ 1.30 = $130K fully-loaded ``` ### Step 5: Project Cash Flow Calculate monthly cash position and runway. **Monthly Cash Flow:** ``` Beginning Cash + Revenue Collected (consider payment terms) - Operating Expenses Paid - CapEx = Ending Cash ``` **Runway Calculation:** ``` If Ending Cash < 0: Funding Need = Negative Cash Balance Runway = 0 Else: Runway = Ending Cash / Average Monthly Burn ``` ### Step 6: Calculate Key Metrics Track metrics that matter for stage. **Revenue Metrics:** - MRR / ARR - Growth rate (MoM, YoY) - Revenue by segment or cohort **Unit Economics:** - CAC (Customer Acquisition Cost) - LTV (Lifetime Value) - CAC Payback Period - LTV / CAC Ratio **Efficiency Metrics:** - Burn multiple (Net Burn / Net New ARR) - Magic number (Net New ARR / S&M Spend) - Rule of 40 (Growth % + Profit Margin %) **Cash Metrics:** - Monthly burn rate - Runway (months) - Cash efficiency ### Step 7: Scenario Analysis Create three scenarios with different assumptions. **Variable Assumptions:** - Customer acquisition rate (±30%) - Churn rate (±20%) - Average contract value (±15%) - CAC (±25%) **Fixed Assumptions:** - Pricing structure - Core operating expenses - Hiring plan (adjust timing, not roles) ## Business Model Templates ### SaaS Financial Model **Revenue Drivers:** - New MRR (customers Ɨ ARPU) - Expansion MRR (upsells) - Contraction MRR (downgrades) - Churned MRR (lost customers) **Key Ratios:** - Gross margin: 75-85% - S&M as % revenue: 40-60% (early stage) - CAC payback: < 12 months - Net retention: 100-120% **Example Projection:** ``` Year 1: $500K ARR, 50 customers, $100K MRR by Dec Year 2: $2.5M ARR, 200 customers, $208K MRR by Dec Year 3: $8M ARR, 600 customers, $667K MRR by Dec ``` ### Marketplace Financial Model **Revenue Drivers:** - GMV (Gross Merchandise Value) - Take rate (% of GMV) - Net revenue = GMV Ɨ Take rate **Key Ratios:** - Take rate: 10-30% depending on category - CAC for buyers vs. sellers - Contribution margin: 60-70% **Example Projection:** ``` Year 1: $5M GMV, 15% take rate = $750K revenue Year 2: $20M GMV, 15% take rate = $3M revenue Year 3: $60M GMV, 15% take rate = $9M revenue ``` ### E-Commerce Financial Model **Revenue Drivers:** - Traffic (visitors) - Conversion rate - Average order value (AOV) - Purchase frequency **Key Ratios:** - Gross margin: 40-60% - Contribution margin: 20-35% - CAC payback: 3-6 months ### Services / Agency Financial Model **Revenue Drivers:** - Billable hours or projects - Hourly rate or project fee - Utilization rate - Team capacity **Key Ratios:** - Gross margin: 50-70% - Utilization: 70-85% - Revenue per employee ## Fundraising Integration ### Funding Scenario Modeling **Pre-Money Valuation:** Based on metrics and comparables. **Dilution:** ``` Post-Money = Pre-Money + Investment Dilution % = Investment / Post-Money ``` **Use of Funds:** Allocate funding to extend runway and achieve milestones. **Example:** ``` Raise: $5M at $20M pre-money Post-Money: $25M Dilution: 20% Use of Funds: - Product Development: $2M (40%) - Sales & Marketing: $2M (40%) - G&A and Operations: $0.5M (10%) - Working Capital: $0.5M (10%) ``` ### Milestone-Based Planning **Identify Key Milestones:** - Product launch - First $1M ARR - Break-even on CAC - Series A fundraise **Funding Amount:** Ensure runway to achieve next milestone + 6 months buffer. ## Common Pitfalls **Pitfall 1: Overly Optimistic Revenue** - New startups rarely hit aggressive projections - Use conservative customer acquisition assumptions - Model realistic churn rates **Pitfall 2: Underestimating Costs** - Add 20% buffer to expense estimates - Include fully-loaded compensation - Account for software and tools **Pitfall 3: Ignoring Cash Flow Timing** - Revenue ≠ cash (payment terms) - Expenses paid before revenue collected - Model cash conversion carefully **Pitfall 4: Static Headcount** - Hiring takes time (3-6 months to fill roles) - Ramp time for productivity (3-6 months) - Account for attrition (10-15% annually) **Pitfall 5: Not Scenario Planning** - Single scenario is never accurate - Always model conservative case - Plan for what you'll do if base case fails ## Model Validation **Sanity Checks:** - [ ] Revenue growth rate is achievable (3x in Year 2, 2x in Year 3) - [ ] Unit economics are realistic (LTV/CAC > 3, payback < 18 months) - [ ] Burn multiple is reasonable (< 2.0 in Year 2-3) - [ ] Headcount scales with revenue (revenue per employee growing) - [ ] Gross margin is appropriate for business model - [ ] S&M spending aligns with CAC and growth targets **Benchmark Against Peers:** Compare key metrics to similar companies at similar stage. **Investor Feedback:** Share model with advisors or investors for feedback on assumptions. ## Additional Resources ### Reference Files For detailed model structures and advanced techniques: - **`references/model-templates.md`** - Complete financial model templates by business model - **`references/unit-economics.md`** - Deep dive on CAC, LTV, payback, and efficiency metrics - **`references/fundraising-scenarios.md`** - Modeling funding rounds and dilution ### Example Files Working financial models with formulas: - **`examples/saas-financial-model.md`** - Complete 3-year SaaS model with cohort analysis - **`examples/marketplace-model.md`** - Marketplace GMV and take rate projections - **`examples/scenario-analysis.md`** - Three-scenario framework with sensitivities ## Quick Start To create a startup financial model: 1. **Define business model** - Revenue drivers and pricing 2. **Project revenue** - Cohort-based with retention 3. **Model costs** - COGS, S&M, R&D, G&A by month 4. **Plan headcount** - Hiring by role and department 5. **Calculate cash flow** - Revenue - expenses = burn/runway 6. **Compute metrics** - CAC, LTV, burn multiple, runway 7. **Create scenarios** - Conservative, base, optimistic 8. **Validate assumptions** - Sanity check and benchmark 9. **Integrate fundraising** - Model funding rounds and milestones For complete templates and formulas, reference the `references/` and `examples/` files.
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

startup-metrics-framework

This skill should be used when the user asks about "key startup

business
⭐1
# Startup Metrics Framework Comprehensive guide to tracking, calculating, and optimizing key performance metrics for different startup business models from seed through Series A. ## Overview Track the right metrics at the right stage. Focus on unit economics, growth efficiency, and cash management metrics that matter for fundraising and operational excellence. ## Universal Startup Metrics ### Revenue Metrics **MRR (Monthly Recurring Revenue)** ``` MRR = Ī£ (Active Subscriptions Ɨ Monthly Price) ``` **ARR (Annual Recurring Revenue)** ``` ARR = MRR Ɨ 12 ``` **Growth Rate** ``` MoM Growth = (This Month MRR - Last Month MRR) / Last Month MRR YoY Growth = (This Year ARR - Last Year ARR) / Last Year ARR ``` **Target Benchmarks:** - Seed stage: 15-20% MoM growth - Series A: 10-15% MoM growth, 3-5x YoY - Series B+: 100%+ YoY (Rule of 40) ### Unit Economics **CAC (Customer Acquisition Cost)** ``` CAC = Total S&M Spend / New Customers Acquired ``` Include: Sales salaries, marketing spend, tools, overhead **LTV (Lifetime Value)** ``` LTV = ARPU Ɨ Gross Margin% Ɨ (1 / Churn Rate) ``` Simplified: ``` LTV = ARPU Ɨ Average Customer Lifetime Ɨ Gross Margin% ``` **LTV:CAC Ratio** ``` LTV:CAC = LTV / CAC ``` **Benchmarks:** - LTV:CAC > 3.0 = Healthy - LTV:CAC 1.0-3.0 = Needs improvement - LTV:CAC < 1.0 = Unsustainable **CAC Payback Period** ``` CAC Payback = CAC / (ARPU Ɨ Gross Margin%) ``` **Benchmarks:** - < 12 months = Excellent - 12-18 months = Good - > 24 months = Concerning ### Cash Efficiency Metrics **Burn Rate** ``` Monthly Burn = Monthly Revenue - Monthly Expenses ``` Negative burn = losing money (typical early-stage) **Runway** ``` Runway (months) = Cash Balance / Monthly Burn Rate ``` **Target:** Always maintain 12-18 months runway **Burn Multiple** ``` Burn Multiple = Net Burn / Net New ARR ``` **Benchmarks:** - < 1.0 = Exceptional efficiency - 1.0-1.5 = Good - 1.5-2.0 = Acceptable - > 2.0 = Inefficient Lower is better (spending less to generate ARR) ## SaaS Metrics ### Revenue Composition **New MRR** New customers Ɨ ARPU **Expansion MRR** Upsells and cross-sells from existing customers **Contraction MRR** Downgrades from existing customers **Churned MRR** Lost customers **Net New MRR Formula:** ``` Net New MRR = New MRR + Expansion MRR - Contraction MRR - Churned MRR ``` ### Retention Metrics **Logo Retention** ``` Logo Retention = (Customers End - New Customers) / Customers Start ``` **Dollar Retention (NDR - Net Dollar Retention)** ``` NDR = (ARR Start + Expansion - Contraction - Churn) / ARR Start ``` **Benchmarks:** - NDR > 120% = Best-in-class - NDR 100-120% = Good - NDR < 100% = Needs work **Gross Retention** ``` Gross Retention = (ARR Start - Churn - Contraction) / ARR Start ``` **Benchmarks:** - > 90% = Excellent - 85-90% = Good - < 85% = Concerning ### SaaS-Specific Metrics **Magic Number** ``` Magic Number = Net New ARR (quarter) / S&M Spend (prior quarter) ``` **Benchmarks:** - > 0.75 = Efficient, ready to scale - 0.5-0.75 = Moderate efficiency - < 0.5 = Inefficient, don't scale yet **Rule of 40** ``` Rule of 40 = Revenue Growth Rate% + Profit Margin% ``` **Benchmarks:** - > 40% = Excellent - 20-40% = Acceptable - < 20% = Needs improvement **Example:** 50% growth + (10%) margin = 40% āœ“ **Quick Ratio** ``` Quick Ratio = (New MRR + Expansion MRR) / (Churned MRR + Contraction MRR) ``` **Benchmarks:** - > 4.0 = Healthy growth - 2.0-4.0 = Moderate - < 2.0 = Churn problem ## Marketplace Metrics ### GMV (Gross Merchandise Value) **Total Transaction Volume:** ``` GMV = Ī£ (Transaction Value) ``` **Growth Rate:** ``` GMV Growth Rate = (Current Period GMV - Prior Period GMV) / Prior Period GMV ``` **Target:** 20%+ MoM early-stage ### Take Rate ``` Take Rate = Net Revenue / GMV ``` **Typical Ranges:** - Payment processors: 2-3% - E-commerce marketplaces: 10-20% - Service marketplaces: 15-25% - High-value B2B: 5-15% ### Marketplace Liquidity **Time to Transaction** How long from listing to sale/match? **Fill Rate** % of requests that result in transaction **Repeat Rate** % of users who transact multiple times **Benchmarks:** - Fill rate > 80% = Strong liquidity - Repeat rate > 60% = Strong retention ### Marketplace Balance **Supply/Demand Ratio:** Track relative growth of supply and demand sides. **Warning Signs:** - Too much supply: Low fill rates, frustrated suppliers - Too much demand: Long wait times, frustrated customers **Goal:** Balanced growth (1:1 ratio ideal, but varies by model) ## Consumer/Mobile Metrics ### Engagement Metrics **DAU (Daily Active Users)** Unique users active each day **MAU (Monthly Active Users)** Unique users active each month **DAU/MAU Ratio** ``` DAU/MAU = DAU / MAU ``` **Benchmarks:** - > 50% = Exceptional (daily habit) - 20-50% = Good - < 20% = Weak engagement **Session Frequency** Average sessions per user per day/week **Session Duration** Average time spent per session ### Retention Curves **Day 1 Retention:** % users who return next day **Day 7 Retention:** % users active 7 days after signup **Day 30 Retention:** % users active 30 days after signup **Benchmarks (Day 30):** - > 40% = Excellent - 25-40% = Good - < 25% = Weak **Retention Curve Shape:** - Flattening curve = good (users becoming habitual) - Steep decline = poor product-market fit ### Viral Coefficient (K-Factor) ``` K-Factor = Invites per User Ɨ Invite Conversion Rate ``` **Example:** 10 invites/user Ɨ 20% conversion = 2.0 K-factor **Benchmarks:** - K > 1.0 = Viral growth - K = 0.5-1.0 = Strong referrals - K < 0.5 = Weak virality ## B2B Metrics ### Sales Efficiency **Win Rate** ``` Win Rate = Deals Won / Total Opportunities ``` **Target:** 20-30% for new sales team, 30-40% mature **Sales Cycle Length** Average days from opportunity to close **Shorter is better:** - SMB: 30-60 days - Mid-market: 60-120 days - Enterprise: 120-270 days **Average Contract Value (ACV)** ``` ACV = Total Contract Value / Contract Length (years) ``` ### Pipeline Metrics **Pipeline Coverage** ``` Pipeline Coverage = Total Pipeline Value / Quota ``` **Target:** 3-5x coverage (3-5x pipeline needed to hit quota) **Conversion Rates by Stage:** - Lead → Opportunity: 10-20% - Opportunity → Demo: 50-70% - Demo → Proposal: 30-50% - Proposal → Close: 20-40% ## Metrics by Stage ### Pre-Seed (Product-Market Fit) **Focus Metrics:** 1. Active users growth 2. User retention (Day 7, Day 30) 3. Core engagement (sessions, features used) 4. Qualitative feedback (NPS, interviews) **Don't worry about:** - Revenue (may be zero) - CAC (not optimizing yet) - Unit economics ### Seed ($500K-$2M ARR) **Focus Metrics:** 1. MRR growth rate (15-20% MoM) 2. CAC and LTV (establish baseline) 3. Gross retention (> 85%) 4. Core product engagement **Start tracking:** - Sales efficiency - Burn rate and runway ### Series A ($2M-$10M ARR) **Focus Metrics:** 1. ARR growth (3-5x YoY) 2. Unit economics (LTV:CAC > 3, payback < 18 months) 3. Net dollar retention (> 100%) 4. Burn multiple (< 2.0) 5. Magic number (> 0.5) **Mature tracking:** - Rule of 40 - Sales efficiency - Pipeline coverage ## Metric Tracking Best Practices ### Data Infrastructure **Requirements:** - Single source of truth (analytics platform) - Real-time or daily updates - Automated calculations - Historical tracking **Tools:** - Mixpanel, Amplitude (product analytics) - ChartMogul, Baremetrics (SaaS metrics) - Looker, Tableau (BI dashboards) ### Reporting Cadence **Daily:** - MRR, active users - Sign-ups, conversions **Weekly:** - Growth rates - Retention cohorts - Sales pipeline **Monthly:** - Full metric suite - Board reporting - Investor updates **Quarterly:** - Trend analysis - Benchmarking - Strategy review ### Common Mistakes **Mistake 1: Vanity Metrics** Don't focus on: - Total users (without retention) - Page views (without engagement) - Downloads (without activation) Focus on actionable metrics tied to value. **Mistake 2: Too Many Metrics** Track 5-7 core metrics intensely, not 50 loosely. **Mistake 3: Ignoring Unit Economics** CAC and LTV are critical even at seed stage. **Mistake 4: Not Segmenting** Break down metrics by customer segment, channel, cohort. **Mistake 5: Gaming Metrics** Optimize for real business outcomes, not dashboard numbers. ## Investor Metrics ### What VCs Want to See **Seed Round:** - MRR growth rate - User retention - Early unit economics - Product engagement **Series A:** - ARR and growth rate - CAC payback < 18 months - LTV:CAC > 3.0 - Net dollar retention > 100% - Burn multiple < 2.0 **Series B+:** - Rule of 40 > 40% - Efficient growth (magic number) - Path to profitability - Market leadership metrics ### Metric Presentation **Dashboard Format:** ``` Current MRR: $250K (↑ 18% MoM) ARR: $3.0M (↑ 280% YoY) CAC: $1,200 | LTV: $4,800 | LTV:CAC = 4.0x NDR: 112% | Logo Retention: 92% Burn: $180K/mo | Runway: 18 months ``` **Include:** - Current value - Growth rate or trend - Context (target, benchmark) ## Additional Resources ### Reference Files - **`references/metric-definitions.md`** - Complete definitions and formulas for 50+ metrics - **`references/benchmarks-by-stage.md`** - Target ranges for each metric by company stage - **`references/calculation-examples.md`** - Step-by-step calculation examples ### Example Files - **`examples/saas-metrics-dashboard.md`** - Complete metrics suite for B2B SaaS company - **`examples/marketplace-metrics.md`** - Marketplace-specific metrics with examples - **`examples/investor-metrics-deck.md`** - How to present metrics for fundraising ## Quick Start To implement startup metrics framework: 1. **Identify business model** - SaaS, marketplace, consumer, B2B 2. **Choose 5-7 core metrics** - Based on stage and model 3. **Establish tracking** - Set up analytics and dashboards 4. **Calculate unit economics** - CAC, LTV, payback 5. **Set targets** - Use benchmarks for goals 6. **Review regularly** - Weekly for core metrics 7. **Share with team** - Align on goals and progress 8. **Update investors** - Monthly/quarterly reporting For detailed definitions, benchmarks, and examples, see `references/` and `examples/`.
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

team-composition-analysis

This skill should be used when the user asks to "plan team

business
⭐1
# Team Composition Analysis Design optimal team structures, hiring plans, compensation strategies, and equity allocation for early-stage startups from pre-seed through Series A. ## Overview Build the right team at the right time with appropriate compensation and equity. Plan role-by-role hiring aligned with revenue milestones, budget constraints, and market benchmarks. ## Team Structure by Stage ### Pre-Seed (0-$500K ARR) **Team Size: 2-5 people** **Core Roles:** - Founders (2-3): Product, engineering, business - First engineer (if needed) - Contract roles: Design, marketing **Focus:** Build and validate product-market fit ### Seed ($500K-$2M ARR) **Team Size: 5-15 people** **Key Hires:** - Engineering lead + 2-3 engineers - First sales/business development - Product manager - Marketing/growth lead **Focus:** Scale product and prove repeatable sales ### Series A ($2M-$10M ARR) **Team Size: 15-50 people** **Department Build-Out:** - Engineering (40%): 6-20 people - Sales & Marketing (30%): 5-15 people - Customer Success (10%): 2-5 people - G&A (10%): 2-5 people - Product (10%): 2-5 people **Focus:** Scale revenue and build repeatable processes ## Role-by-Role Planning ### Engineering Team **Pre-Seed:** - Founders write code - 0-1 contract developers **Seed:** - Engineering Lead (first $150K-$180K) - 2-3 Full-Stack Engineers ($120K-$150K) - 1 Frontend or Backend Specialist ($130K-$160K) **Series A:** - VP Engineering ($180K-$250K + equity) - 2-3 Senior Engineers ($150K-$180K) - 3-5 Mid-Level Engineers ($120K-$150K) - 1-2 Junior Engineers ($90K-$120K) - 1 DevOps/Infrastructure ($140K-$170K) ### Sales & Marketing **Pre-Seed:** - Founders do sales - Contract marketing help **Seed:** - First Sales Hire / Head of Sales ($120K-$150K + commission) - Marketing/Growth Lead ($100K-$140K) - SDR or BDR (if B2B) ($50K-$70K + commission) **Series A:** - VP Sales ($150K-$200K + commission + equity) - 3-5 Account Executives ($80K-$120K + commission) - 2-3 SDRs/BDRs ($50K-$70K + commission) - Marketing Manager ($90K-$130K) - Content/Demand Gen ($70K-$100K) ### Product Team **Pre-Seed:** - Founder as product lead **Seed:** - First Product Manager ($120K-$150K) - Contract designer **Series A:** - Head of Product ($150K-$180K) - 1-2 Product Managers ($120K-$150K) - Product Designer ($100K-$140K) - UX Researcher (optional) ($90K-$130K) ### Customer Success **Pre-Seed:** - Founders handle support **Seed:** - First CS hire (optional) ($60K-$90K) **Series A:** - CS Manager ($100K-$130K) - 2-4 CS Representatives ($60K-$90K) - Support Engineer (technical) ($80K-$120K) ### G&A (General & Administrative) **Pre-Seed:** - Contractors (accounting, legal) **Seed:** - Operations/Office Manager ($70K-$100K) - Contract CFO **Series A:** - CFO or Finance Lead ($150K-$200K) - Recruiter ($80K-$120K) - Office Manager / EA ($60K-$90K) ## Compensation Strategy ### Base Salary Benchmarks (US, 2024) **Engineering:** - Junior: $90K-$120K - Mid-Level: $120K-$150K - Senior: $150K-$180K - Staff/Principal: $180K-$220K - Engineering Manager: $160K-$200K - VP Engineering: $180K-$250K **Sales:** - SDR/BDR: $50K-$70K base + $50K-$70K commission - Account Executive: $80K-$120K base + $80K-$120K commission - Sales Manager: $120K-$160K base + $80K-$120K commission - VP Sales: $150K-$200K base + $150K-$200K commission **Product:** - Product Manager: $120K-$150K - Senior PM: $150K-$180K - Head of Product: $150K-$180K - VP Product: $180K-$220K **Marketing:** - Marketing Manager: $90K-$130K - Content/Demand Gen: $70K-$100K - Head of Marketing: $130K-$170K - VP Marketing: $150K-$200K **Customer Success:** - CS Representative: $60K-$90K - CS Manager: $100K-$130K - VP Customer Success: $140K-$180K ### Total Compensation Formula ``` Total Comp = Base Salary Ɨ 1.30 (benefits & taxes) + Equity Value ``` **Fully-Loaded Cost:** - Base salary - Payroll taxes (7.65% FICA) - Benefits (health insurance, 401k): $10K-$15K per employee - Other (workspace, equipment, software): $5K-$10K per employee **Rule of Thumb:** Multiply base salary by 1.3-1.4 for fully-loaded cost ### Geographic Adjustments **San Francisco / New York:** +20-30% above benchmarks **Seattle / Boston / Los Angeles:** +10-20% **Austin / Denver / Chicago:** +0-10% **Remote / Other US Cities:** -10-20% **International:** Varies widely by country ## Equity Allocation ### Equity by Role and Stage **Founders:** - First founder: 40-60% - Second founder: 20-40% - Third founder: 10-20% - Vesting: 4 years with 1-year cliff **Early Employees (Pre-Seed):** - First engineer: 0.5-2.0% - First 5 employees: 0.25-1.0% each **Seed Stage Hires:** - VP/Head level: 0.5-1.5% - Senior IC: 0.1-0.5% - Mid-level: 0.05-0.25% - Junior: 0.01-0.1% **Series A Hires:** - C-level (CTO, CFO): 1.0-3.0% - VP level: 0.3-1.0% - Director level: 0.1-0.5% - Senior IC: 0.05-0.2% - Mid-level: 0.01-0.1% - Junior: 0.005-0.05% ### Equity Pool Sizing **Option Pool by Round:** - Pre-Seed: 10-15% reserved - Seed: 10-15% top-up - Series A: 10-15% top-up - Series B+: 5-10% per round **Pre-Funding Dilution:** Investors often require option pool creation before investment, diluting founders. **Example:** ``` Pre-money: $10M Investors want 15% option pool post-money Calculation: Post-money: $15M ($10M + $5M investment) Option pool: $2.25M (15% Ɨ $15M) Founders diluted by pool creation before new money ``` ## Organizational Design ### Reporting Structure **Pre-Seed:** ``` Founders (flat structure) ā”œā”€ā”€ Contractors └── First hires (report to founders) ``` **Seed:** ``` CEO ā”œā”€ā”€ Engineering Lead (2-4 engineers) ā”œā”€ā”€ Sales/Growth Lead (1-2 reps) ā”œā”€ā”€ Product Manager └── Operations ``` **Series A:** ``` CEO ā”œā”€ā”€ CTO / VP Engineering (6-20 people) │ ā”œā”€ā”€ Engineering Manager(s) │ └── Individual Contributors ā”œā”€ā”€ VP Sales (5-15 people) │ ā”œā”€ā”€ Sales Manager │ ā”œā”€ā”€ Account Executives │ └── SDRs ā”œā”€ā”€ Head of Product (2-5 people) │ ā”œā”€ā”€ Product Managers │ └── Designers ā”œā”€ā”€ Head of Customer Success (2-5 people) └── CFO / Finance Lead (2-5 people) ā”œā”€ā”€ Recruiter └── Operations ``` ### Span of Control **Manager Ratios:** - First-line managers: 4-8 direct reports - Directors: 3-5 direct reports (managers) - VPs: 3-5 direct reports (directors) - CEO: 5-8 direct reports (executive team) ## Full-Time vs. Contract ### Use Full-Time for: - Core product development - Sales (revenue-generating roles) - Mission-critical operations - Institutional knowledge roles ### Use Contractors for: - Specialized short-term needs (legal, accounting) - Variable workload (design, marketing campaigns) - Skills outside core competency - Testing role before FTE hire - Geographic expansion before permanent presence ### Cost Comparison **Full-Time:** - Lower hourly cost - Benefits and overhead - Long-term commitment - Cultural fit matters **Contract:** - Higher hourly rate ($75-$200/hour vs. $40-$100/hour FTE equivalent) - No benefits or overhead - Flexible engagement - Easier to scale up/down ## Hiring Velocity ### Realistic Timeline **Role Opening to Hire:** - Junior: 6-8 weeks - Mid-Level: 8-12 weeks - Senior: 12-16 weeks - Executive: 16-24 weeks **Time to Productivity:** - Junior: 4-6 months - Mid-Level: 2-4 months - Senior: 1-3 months - Executive: 3-6 months ### Planning Buffer Always add 2-3 months buffer to hiring plans. **Example:** If need engineer by July 1: - Start recruiting: April 1 (12 weeks) - Productivity: September 1 (2 months ramp) ## Budget Planning ### Compensation as % of Revenue **Early Stage (Seed):** - Total comp: 120-150% of revenue (burning cash to grow) - Engineering: 50-60% - Sales: 30-40% - Other: 20-30% **Growth Stage (Series A):** - Total comp: 70-100% of revenue - Engineering: 35-45% - Sales: 25-35% - Other: 20-30% ### Headcount Budget Formula ``` Total Comp Budget = Ī£ (Role Count Ɨ Fully-Loaded Cost Ɨ % of Year) Example: 3 Engineers Ɨ $202K Ɨ 100% = $606K 2 AEs Ɨ $230K Ɨ 75% (mid-year start) = $345K 1 PM Ɨ $162K Ɨ 100% = $162K Total: $1.1M ``` ## Additional Resources ### Reference Files - **`references/compensation-benchmarks.md`** - Detailed salary data by role, level, and location - **`references/equity-calculator.md`** - Equity sizing formulas and dilution scenarios ### Example Files - **`examples/seed-stage-hiring-plan.md`** - Complete hiring plan for seed-stage SaaS company - **`examples/org-chart-evolution.md`** - Organizational design from 5 to 50 people ## Quick Start To plan team composition: 1. **Identify stage** - Pre-seed, seed, or Series A 2. **Define roles** - What functions are needed now 3. **Prioritize hires** - Critical path for business goals 4. **Set compensation** - Base salary + equity by level 5. **Plan timeline** - Account for recruiting and ramp time 6. **Calculate budget** - Fully-loaded cost Ɨ headcount 7. **Design org chart** - Reporting structure and span of control 8. **Allocate equity** - Fair allocation that preserves pool For detailed compensation benchmarks and hiring plan templates, see `references/` and `examples/`.
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

code-review-excellence

Master effective code review practices to provide constructive

coding
⭐1
# Code Review Excellence Transform code reviews from gatekeeping to knowledge sharing through constructive feedback, systematic analysis, and collaborative improvement. ## When to Use This Skill - Reviewing pull requests and code changes - Establishing code review standards for teams - Mentoring junior developers through reviews - Conducting architecture reviews - Creating review checklists and guidelines - Improving team collaboration - Reducing code review cycle time - Maintaining code quality standards ## Core Principles ### 1. The Review Mindset **Goals of Code Review:** - Catch bugs and edge cases - Ensure code maintainability - Share knowledge across team - Enforce coding standards - Improve design and architecture - Build team culture **Not the Goals:** - Show off knowledge - Nitpick formatting (use linters) - Block progress unnecessarily - Rewrite to your preference ### 2. Effective Feedback **Good Feedback is:** - Specific and actionable - Educational, not judgmental - Focused on the code, not the person - Balanced (praise good work too) - Prioritized (critical vs nice-to-have) ```markdown āŒ Bad: "This is wrong." āœ… Good: "This could cause a race condition when multiple users access simultaneously. Consider using a mutex here." āŒ Bad: "Why didn't you use X pattern?" āœ… Good: "Have you considered the Repository pattern? It would make this easier to test. Here's an example: [link]" āŒ Bad: "Rename this variable." āœ… Good: "[nit] Consider `userCount` instead of `uc` for clarity. Not blocking if you prefer to keep it." ``` ### 3. Review Scope **What to Review:** - Logic correctness and edge cases - Security vulnerabilities - Performance implications - Test coverage and quality - Error handling - Documentation and comments - API design and naming - Architectural fit **What Not to Review Manually:** - Code formatting (use Prettier, Black, etc.) - Import organization - Linting violations - Simple typos ## Review Process ### Phase 1: Context Gathering (2-3 minutes) ```markdown Before diving into code, understand: 1. Read PR description and linked issue 2. Check PR size (>400 lines? Ask to split) 3. Review CI/CD status (tests passing?) 4. Understand the business requirement 5. Note any relevant architectural decisions ``` ### Phase 2: High-Level Review (5-10 minutes) ```markdown 1. **Architecture & Design** - Does the solution fit the problem? - Are there simpler approaches? - Is it consistent with existing patterns? - Will it scale? 2. **File Organization** - Are new files in the right places? - Is code grouped logically? - Are there duplicate files? 3. **Testing Strategy** - Are there tests? - Do tests cover edge cases? - Are tests readable? ``` ### Phase 3: Line-by-Line Review (10-20 minutes) ```markdown For each file: 1. **Logic & Correctness** - Edge cases handled? - Off-by-one errors? - Null/undefined checks? - Race conditions? 2. **Security** - Input validation? - SQL injection risks? - XSS vulnerabilities? - Sensitive data exposure? 3. **Performance** - N+1 queries? - Unnecessary loops? - Memory leaks? - Blocking operations? 4. **Maintainability** - Clear variable names? - Functions doing one thing? - Complex code commented? - Magic numbers extracted? ``` ### Phase 4: Summary & Decision (2-3 minutes) ```markdown 1. Summarize key concerns 2. Highlight what you liked 3. Make clear decision: - āœ… Approve - šŸ’¬ Comment (minor suggestions) - šŸ”„ Request Changes (must address) 4. Offer to pair if complex ``` ## Review Techniques ### Technique 1: The Checklist Method ```markdown ## Security Checklist - [ ] User input validated and sanitized - [ ] SQL queries use parameterization - [ ] Authentication/authorization checked - [ ] Secrets not hardcoded - [ ] Error messages don't leak info ## Performance Checklist - [ ] No N+1 queries - [ ] Database queries indexed - [ ] Large lists paginated - [ ] Expensive operations cached - [ ] No blocking I/O in hot paths ## Testing Checklist - [ ] Happy path tested - [ ] Edge cases covered - [ ] Error cases tested - [ ] Test names are descriptive - [ ] Tests are deterministic ``` ### Technique 2: The Question Approach Instead of stating problems, ask questions to encourage thinking: ```markdown āŒ "This will fail if the list is empty." āœ… "What happens if `items` is an empty array?" āŒ "You need error handling here." āœ… "How should this behave if the API call fails?" āŒ "This is inefficient." āœ… "I see this loops through all users. Have we considered the performance impact with 100k users?" ``` ### Technique 3: Suggest, Don't Command ````markdown ## Use Collaborative Language āŒ "You must change this to use async/await" āœ… "Suggestion: async/await might make this more readable: `typescript async function fetchUser(id: string) { const user = await db.query('SELECT * FROM users WHERE id = ?', id); return user; } ` What do you think?" āŒ "Extract this into a function" āœ… "This logic appears in 3 places. Would it make sense to extract it into a shared utility function?" ```` ### Technique 4: Differentiate Severity ```markdown Use labels to indicate priority: šŸ”“ [blocking] - Must fix before merge 🟔 [important] - Should fix, discuss if disagree 🟢 [nit] - Nice to have, not blocking šŸ’” [suggestion] - Alternative approach to consider šŸ“š [learning] - Educational comment, no action needed šŸŽ‰ [praise] - Good work, keep it up! Example: "šŸ”“ [blocking] This SQL query is vulnerable to injection. Please use parameterized queries." "🟢 [nit] Consider renaming `data` to `userData` for clarity." "šŸŽ‰ [praise] Excellent test coverage! This will catch edge cases." ``` ## Language-Specific Patterns ### Python Code Review ```python # Check for Python-specific issues # āŒ Mutable default arguments def add_item(item, items=[]): # Bug! Shared across calls items.append(item) return items # āœ… Use None as default def add_item(item, items=None): if items is None: items = [] items.append(item) return items # āŒ Catching too broad try: result = risky_operation() except: # Catches everything, even KeyboardInterrupt! pass # āœ… Catch specific exceptions try: result = risky_operation() except ValueError as e: logger.error(f"Invalid value: {e}") raise # āŒ Using mutable class attributes class User: permissions = [] # Shared across all instances! # āœ… Initialize in __init__ class User: def __init__(self): self.permissions = [] ``` ### TypeScript/JavaScript Code Review ```typescript // Check for TypeScript-specific issues // āŒ Using any defeats type safety function processData(data: any) { // Avoid any return data.value; } // āœ… Use proper types interface DataPayload { value: string; } function processData(data: DataPayload) { return data.value; } // āŒ Not handling async errors async function fetchUser(id: string) { const response = await fetch(`/api/users/${id}`); return response.json(); // What if network fails? } // āœ… Handle errors properly async function fetchUser(id: string): Promise<User> { try { const response = await fetch(`/api/users/${id}`); if (!response.ok) { throw new Error(`HTTP ${response.status}`); } return await response.json(); } catch (error) { console.error('Failed to fetch user:', error); throw error; } } // āŒ Mutation of props function UserProfile({ user }: Props) { user.lastViewed = new Date(); // Mutating prop! return <div>{user.name}</div>; } // āœ… Don't mutate props function UserProfile({ user, onView }: Props) { useEffect(() => { onView(user.id); // Notify parent to update }, [user.id]); return <div>{user.name}</div>; } ``` ## Advanced Review Patterns ### Pattern 1: Architectural Review ```markdown When reviewing significant changes: 1. **Design Document First** - For large features, request design doc before code - Review design with team before implementation - Agree on approach to avoid rework 2. **Review in Stages** - First PR: Core abstractions and interfaces - Second PR: Implementation - Third PR: Integration and tests - Easier to review, faster to iterate 3. **Consider Alternatives** - "Have we considered using [pattern/library]?" - "What's the tradeoff vs. the simpler approach?" - "How will this evolve as requirements change?" ``` ### Pattern 2: Test Quality Review ```typescript // āŒ Poor test: Implementation detail testing test('increments counter variable', () => { const component = render(<Counter />); const button = component.getByRole('button'); fireEvent.click(button); expect(component.state.counter).toBe(1); // Testing internal state }); // āœ… Good test: Behavior testing test('displays incremented count when clicked', () => { render(<Counter />); const button = screen.getByRole('button', { name: /increment/i }); fireEvent.click(button); expect(screen.getByText('Count: 1')).toBeInTheDocument(); }); // Review questions for tests: // - Do tests describe behavior, not implementation? // - Are test names clear and descriptive? // - Do tests cover edge cases? // - Are tests independent (no shared state)? // - Can tests run in any order? ``` ### Pattern 3: Security Review ```markdown ## Security Review Checklist ### Authentication & Authorization - [ ] Is authentication required where needed? - [ ] Are authorization checks before every action? - [ ] Is JWT validation proper (signature, expiry)? - [ ] Are API keys/secrets properly secured? ### Input Validation - [ ] All user inputs validated? - [ ] File uploads restricted (size, type)? - [ ] SQL queries parameterized? - [ ] XSS protection (escape output)? ### Data Protection - [ ] Passwords hashed (bcrypt/argon2)? - [ ] Sensitive data encrypted at rest? - [ ] HTTPS enforced for sensitive data? - [ ] PII handled according to regulations? ### Common Vulnerabilities - [ ] No eval() or similar dynamic execution? - [ ] No hardcoded secrets? - [ ] CSRF protection for state-changing operations? - [ ] Rate limiting on public endpoints? ``` ## Giving Difficult Feedback ### Pattern: The Sandwich Method (Modified) ```markdown Traditional: Praise + Criticism + Praise (feels fake) Better: Context + Specific Issue + Helpful Solution Example: "I noticed the payment processing logic is inline in the controller. This makes it harder to test and reuse. [Specific Issue] The calculateTotal() function mixes tax calculation, discount logic, and database queries, making it difficult to unit test and reason about. [Helpful Solution] Could we extract this into a PaymentService class? That would make it testable and reusable. I can pair with you on this if helpful." ``` ### Handling Disagreements ```markdown When author disagrees with your feedback: 1. **Seek to Understand** "Help me understand your approach. What led you to choose this pattern?" 2. **Acknowledge Valid Points** "That's a good point about X. I hadn't considered that." 3. **Provide Data** "I'm concerned about performance. Can we add a benchmark to validate the approach?" 4. **Escalate if Needed** "Let's get [architect/senior dev] to weigh in on this." 5. **Know When to Let Go** If it's working and not a critical issue, approve it. Perfection is the enemy of progress. ``` ## Best Practices 1. **Review Promptly**: Within 24 hours, ideally same day 2. **Limit PR Size**: 200-400 lines max for effective review 3. **Review in Time Blocks**: 60 minutes max, take breaks 4. **Use Review Tools**: GitHub, GitLab, or dedicated tools 5. **Automate What You Can**: Linters, formatters, security scans 6. **Build Rapport**: Emoji, praise, and empathy matter 7. **Be Available**: Offer to pair on complex issues 8. **Learn from Others**: Review others' review comments ## Common Pitfalls - **Perfectionism**: Blocking PRs for minor style preferences - **Scope Creep**: "While you're at it, can you also..." - **Inconsistency**: Different standards for different people - **Delayed Reviews**: Letting PRs sit for days - **Ghosting**: Requesting changes then disappearing - **Rubber Stamping**: Approving without actually reviewing - **Bike Shedding**: Debating trivial details extensively ## Templates ### PR Review Comment Template ```markdown ## Summary [Brief overview of what was reviewed] ## Strengths - [What was done well] - [Good patterns or approaches] ## Required Changes šŸ”“ [Blocking issue 1] šŸ”“ [Blocking issue 2] ## Suggestions šŸ’” [Improvement 1] šŸ’” [Improvement 2] ## Questions ā“ [Clarification needed on X] ā“ [Alternative approach consideration] ## Verdict āœ… Approve after addressing required changes ``` ## Resources - **references/code-review-best-practices.md**: Comprehensive review guidelines - **references/common-bugs-checklist.md**: Language-specific bugs to watch for - **references/security-review-guide.md**: Security-focused review checklist - **assets/pr-review-template.md**: Standard review comment template - **assets/review-checklist.md**: Quick reference checklist - **scripts/pr-analyzer.py**: Analyze PR complexity and suggest reviewers
šŸ‘0
šŸ‘ļø3
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

python-performance-optimization

Profile and optimize Python code using cProfile, memory profilers,

coding
⭐1
# Python Performance Optimization Comprehensive guide to profiling, analyzing, and optimizing Python code for better performance, including CPU profiling, memory optimization, and implementation best practices. ## When to Use This Skill - Identifying performance bottlenecks in Python applications - Reducing application latency and response times - Optimizing CPU-intensive operations - Reducing memory consumption and memory leaks - Improving database query performance - Optimizing I/O operations - Speeding up data processing pipelines - Implementing high-performance algorithms - Profiling production applications ## Core Concepts ### 1. Profiling Types - **CPU Profiling**: Identify time-consuming functions - **Memory Profiling**: Track memory allocation and leaks - **Line Profiling**: Profile at line-by-line granularity - **Call Graph**: Visualize function call relationships ### 2. Performance Metrics - **Execution Time**: How long operations take - **Memory Usage**: Peak and average memory consumption - **CPU Utilization**: Processor usage patterns - **I/O Wait**: Time spent on I/O operations ### 3. Optimization Strategies - **Algorithmic**: Better algorithms and data structures - **Implementation**: More efficient code patterns - **Parallelization**: Multi-threading/processing - **Caching**: Avoid redundant computation - **Native Extensions**: C/Rust for critical paths ## Quick Start ### Basic Timing ```python import time def measure_time(): """Simple timing measurement.""" start = time.time() # Your code here result = sum(range(1000000)) elapsed = time.time() - start print(f"Execution time: {elapsed:.4f} seconds") return result # Better: use timeit for accurate measurements import timeit execution_time = timeit.timeit( "sum(range(1000000))", number=100 ) print(f"Average time: {execution_time/100:.6f} seconds") ``` ## Profiling Tools ### Pattern 1: cProfile - CPU Profiling ```python import cProfile import pstats from pstats import SortKey def slow_function(): """Function to profile.""" total = 0 for i in range(1000000): total += i return total def another_function(): """Another function.""" return [i**2 for i in range(100000)] def main(): """Main function to profile.""" result1 = slow_function() result2 = another_function() return result1, result2 # Profile the code if __name__ == "__main__": profiler = cProfile.Profile() profiler.enable() main() profiler.disable() # Print stats stats = pstats.Stats(profiler) stats.sort_stats(SortKey.CUMULATIVE) stats.print_stats(10) # Top 10 functions # Save to file for later analysis stats.dump_stats("profile_output.prof") ``` **Command-line profiling:** ```bash # Profile a script python -m cProfile -o output.prof script.py # View results python -m pstats output.prof # In pstats: # sort cumtime # stats 10 ``` ### Pattern 2: line_profiler - Line-by-Line Profiling ```python # Install: pip install line-profiler # Add @profile decorator (line_profiler provides this) @profile def process_data(data): """Process data with line profiling.""" result = [] for item in data: processed = item * 2 result.append(processed) return result # Run with: # kernprof -l -v script.py ``` **Manual line profiling:** ```python from line_profiler import LineProfiler def process_data(data): """Function to profile.""" result = [] for item in data: processed = item * 2 result.append(processed) return result if __name__ == "__main__": lp = LineProfiler() lp.add_function(process_data) data = list(range(100000)) lp_wrapper = lp(process_data) lp_wrapper(data) lp.print_stats() ``` ### Pattern 3: memory_profiler - Memory Usage ```python # Install: pip install memory-profiler from memory_profiler import profile @profile def memory_intensive(): """Function that uses lots of memory.""" # Create large list big_list = [i for i in range(1000000)] # Create large dict big_dict = {i: i**2 for i in range(100000)} # Process data result = sum(big_list) return result if __name__ == "__main__": memory_intensive() # Run with: # python -m memory_profiler script.py ``` ### Pattern 4: py-spy - Production Profiling ```bash # Install: pip install py-spy # Profile a running Python process py-spy top --pid 12345 # Generate flamegraph py-spy record -o profile.svg --pid 12345 # Profile a script py-spy record -o profile.svg -- python script.py # Dump current call stack py-spy dump --pid 12345 ``` ## Optimization Patterns ### Pattern 5: List Comprehensions vs Loops ```python import timeit # Slow: Traditional loop def slow_squares(n): """Create list of squares using loop.""" result = [] for i in range(n): result.append(i**2) return result # Fast: List comprehension def fast_squares(n): """Create list of squares using comprehension.""" return [i**2 for i in range(n)] # Benchmark n = 100000 slow_time = timeit.timeit(lambda: slow_squares(n), number=100) fast_time = timeit.timeit(lambda: fast_squares(n), number=100) print(f"Loop: {slow_time:.4f}s") print(f"Comprehension: {fast_time:.4f}s") print(f"Speedup: {slow_time/fast_time:.2f}x") # Even faster for simple operations: map def faster_squares(n): """Use map for even better performance.""" return list(map(lambda x: x**2, range(n))) ``` ### Pattern 6: Generator Expressions for Memory ```python import sys def list_approach(): """Memory-intensive list.""" data = [i**2 for i in range(1000000)] return sum(data) def generator_approach(): """Memory-efficient generator.""" data = (i**2 for i in range(1000000)) return sum(data) # Memory comparison list_data = [i for i in range(1000000)] gen_data = (i for i in range(1000000)) print(f"List size: {sys.getsizeof(list_data)} bytes") print(f"Generator size: {sys.getsizeof(gen_data)} bytes") # Generators use constant memory regardless of size ``` ### Pattern 7: String Concatenation ```python import timeit def slow_concat(items): """Slow string concatenation.""" result = "" for item in items: result += str(item) return result def fast_concat(items): """Fast string concatenation with join.""" return "".join(str(item) for item in items) def faster_concat(items): """Even faster with list.""" parts = [str(item) for item in items] return "".join(parts) items = list(range(10000)) # Benchmark slow = timeit.timeit(lambda: slow_concat(items), number=100) fast = timeit.timeit(lambda: fast_concat(items), number=100) faster = timeit.timeit(lambda: faster_concat(items), number=100) print(f"Concatenation (+): {slow:.4f}s") print(f"Join (generator): {fast:.4f}s") print(f"Join (list): {faster:.4f}s") ``` ### Pattern 8: Dictionary Lookups vs List Searches ```python import timeit # Create test data size = 10000 items = list(range(size)) lookup_dict = {i: i for i in range(size)} def list_search(items, target): """O(n) search in list.""" return target in items def dict_search(lookup_dict, target): """O(1) search in dict.""" return target in lookup_dict target = size - 1 # Worst case for list # Benchmark list_time = timeit.timeit( lambda: list_search(items, target), number=1000 ) dict_time = timeit.timeit( lambda: dict_search(lookup_dict, target), number=1000 ) print(f"List search: {list_time:.6f}s") print(f"Dict search: {dict_time:.6f}s") print(f"Speedup: {list_time/dict_time:.0f}x") ``` ### Pattern 9: Local Variable Access ```python import timeit # Global variable (slow) GLOBAL_VALUE = 100 def use_global(): """Access global variable.""" total = 0 for i in range(10000): total += GLOBAL_VALUE return total def use_local(): """Use local variable.""" local_value = 100 total = 0 for i in range(10000): total += local_value return total # Local is faster global_time = timeit.timeit(use_global, number=1000) local_time = timeit.timeit(use_local, number=1000) print(f"Global access: {global_time:.4f}s") print(f"Local access: {local_time:.4f}s") print(f"Speedup: {global_time/local_time:.2f}x") ``` ### Pattern 10: Function Call Overhead ```python import timeit def calculate_inline(): """Inline calculation.""" total = 0 for i in range(10000): total += i * 2 + 1 return total def helper_function(x): """Helper function.""" return x * 2 + 1 def calculate_with_function(): """Calculation with function calls.""" total = 0 for i in range(10000): total += helper_function(i) return total # Inline is faster due to no call overhead inline_time = timeit.timeit(calculate_inline, number=1000) function_time = timeit.timeit(calculate_with_function, number=1000) print(f"Inline: {inline_time:.4f}s") print(f"Function calls: {function_time:.4f}s") ``` ## Advanced Optimization ### Pattern 11: NumPy for Numerical Operations ```python import timeit import numpy as np def python_sum(n): """Sum using pure Python.""" return sum(range(n)) def numpy_sum(n): """Sum using NumPy.""" return np.arange(n).sum() n = 1000000 python_time = timeit.timeit(lambda: python_sum(n), number=100) numpy_time = timeit.timeit(lambda: numpy_sum(n), number=100) print(f"Python: {python_time:.4f}s") print(f"NumPy: {numpy_time:.4f}s") print(f"Speedup: {python_time/numpy_time:.2f}x") # Vectorized operations def python_multiply(): """Element-wise multiplication in Python.""" a = list(range(100000)) b = list(range(100000)) return [x * y for x, y in zip(a, b)] def numpy_multiply(): """Vectorized multiplication in NumPy.""" a = np.arange(100000) b = np.arange(100000) return a * b py_time = timeit.timeit(python_multiply, number=100) np_time = timeit.timeit(numpy_multiply, number=100) print(f"\nPython multiply: {py_time:.4f}s") print(f"NumPy multiply: {np_time:.4f}s") print(f"Speedup: {py_time/np_time:.2f}x") ``` ### Pattern 12: Caching with functools.lru_cache ```python from functools import lru_cache import timeit def fibonacci_slow(n): """Recursive fibonacci without caching.""" if n < 2: return n return fibonacci_slow(n-1) + fibonacci_slow(n-2) @lru_cache(maxsize=None) def fibonacci_fast(n): """Recursive fibonacci with caching.""" if n < 2: return n return fibonacci_fast(n-1) + fibonacci_fast(n-2) # Massive speedup for recursive algorithms n = 30 slow_time = timeit.timeit(lambda: fibonacci_slow(n), number=1) fast_time = timeit.timeit(lambda: fibonacci_fast(n), number=1000) print(f"Without cache (1 run): {slow_time:.4f}s") print(f"With cache (1000 runs): {fast_time:.4f}s") # Cache info print(f"Cache info: {fibonacci_fast.cache_info()}") ``` ### Pattern 13: Using **slots** for Memory ```python import sys class RegularClass: """Regular class with __dict__.""" def __init__(self, x, y, z): self.x = x self.y = y self.z = z class SlottedClass: """Class with __slots__ for memory efficiency.""" __slots__ = ['x', 'y', 'z'] def __init__(self, x, y, z): self.x = x self.y = y self.z = z # Memory comparison regular = RegularClass(1, 2, 3) slotted = SlottedClass(1, 2, 3) print(f"Regular class size: {sys.getsizeof(regular)} bytes") print(f"Slotted class size: {sys.getsizeof(slotted)} bytes") # Significant savings with many instances regular_objects = [RegularClass(i, i+1, i+2) for i in range(10000)] slotted_objects = [SlottedClass(i, i+1, i+2) for i in range(10000)] print(f"\nMemory for 10000 regular objects: ~{sys.getsizeof(regular) * 10000} bytes") print(f"Memory for 10000 slotted objects: ~{sys.getsizeof(slotted) * 10000} bytes") ``` ### Pattern 14: Multiprocessing for CPU-Bound Tasks ```python import multiprocessing as mp import time def cpu_intensive_task(n): """CPU-intensive calculation.""" return sum(i**2 for i in range(n)) def sequential_processing(): """Process tasks sequentially.""" start = time.time() results = [cpu_intensive_task(1000000) for _ in range(4)] elapsed = time.time() - start return elapsed, results def parallel_processing(): """Process tasks in parallel.""" start = time.time() with mp.Pool(processes=4) as pool: results = pool.map(cpu_intensive_task, [1000000] * 4) elapsed = time.time() - start return elapsed, results if __name__ == "__main__": seq_time, seq_results = sequential_processing() par_time, par_results = parallel_processing() print(f"Sequential: {seq_time:.2f}s") print(f"Parallel: {par_time:.2f}s") print(f"Speedup: {seq_time/par_time:.2f}x") ``` ### Pattern 15: Async I/O for I/O-Bound Tasks ```python import asyncio import aiohttp import time import requests urls = [ "https://httpbin.org/delay/1", "https://httpbin.org/delay/1", "https://httpbin.org/delay/1", "https://httpbin.org/delay/1", ] def synchronous_requests(): """Synchronous HTTP requests.""" start = time.time() results = [] for url in urls: response = requests.get(url) results.append(response.status_code) elapsed = time.time() - start return elapsed, results async def async_fetch(session, url): """Async HTTP request.""" async with session.get(url) as response: return response.status async def asynchronous_requests(): """Asynchronous HTTP requests.""" start = time.time() async with aiohttp.ClientSession() as session: tasks = [async_fetch(session, url) for url in urls] results = await asyncio.gather(*tasks) elapsed = time.time() - start return elapsed, results # Async is much faster for I/O-bound work sync_time, sync_results = synchronous_requests() async_time, async_results = asyncio.run(asynchronous_requests()) print(f"Synchronous: {sync_time:.2f}s") print(f"Asynchronous: {async_time:.2f}s") print(f"Speedup: {sync_time/async_time:.2f}x") ``` ## Database Optimization ### Pattern 16: Batch Database Operations ```python import sqlite3 import time def create_db(): """Create test database.""" conn = sqlite3.connect(":memory:") conn.execute("CREATE TABLE users (id INTEGER PRIMARY KEY, name TEXT)") return conn def slow_inserts(conn, count): """Insert records one at a time.""" start = time.time() cursor = conn.cursor() for i in range(count): cursor.execute("INSERT INTO users (name) VALUES (?)", (f"User {i}",)) conn.commit() # Commit each insert elapsed = time.time() - start return elapsed def fast_inserts(conn, count): """Batch insert with single commit.""" start = time.time() cursor = conn.cursor() data = [(f"User {i}",) for i in range(count)] cursor.executemany("INSERT INTO users (name) VALUES (?)", data) conn.commit() # Single commit elapsed = time.time() - start return elapsed # Benchmark conn1 = create_db() slow_time = slow_inserts(conn1, 1000) conn2 = create_db() fast_time = fast_inserts(conn2, 1000) print(f"Individual inserts: {slow_time:.4f}s") print(f"Batch insert: {fast_time:.4f}s") print(f"Speedup: {slow_time/fast_time:.2f}x") ``` ### Pattern 17: Query Optimization ```python # Use indexes for frequently queried columns """ -- Slow: No index SELECT * FROM users WHERE email = 'user@example.com'; -- Fast: With index CREATE INDEX idx_users_email ON users(email); SELECT * FROM users WHERE email = 'user@example.com'; """ # Use query planning import sqlite3 conn = sqlite3.connect("example.db") cursor = conn.cursor() # Analyze query performance cursor.execute("EXPLAIN QUERY PLAN SELECT * FROM users WHERE email = ?", ("test@example.com",)) print(cursor.fetchall()) # Use SELECT only needed columns # Slow: SELECT * # Fast: SELECT id, name ``` ## Memory Optimization ### Pattern 18: Detecting Memory Leaks ```python import tracemalloc import gc def memory_leak_example(): """Example that leaks memory.""" leaked_objects = [] for i in range(100000): # Objects added but never removed leaked_objects.append([i] * 100) # In real code, this would be an unintended reference def track_memory_usage(): """Track memory allocations.""" tracemalloc.start() # Take snapshot before snapshot1 = tracemalloc.take_snapshot() # Run code memory_leak_example() # Take snapshot after snapshot2 = tracemalloc.take_snapshot() # Compare top_stats = snapshot2.compare_to(snapshot1, 'lineno') print("Top 10 memory allocations:") for stat in top_stats[:10]: print(stat) tracemalloc.stop() # Monitor memory track_memory_usage() # Force garbage collection gc.collect() ``` ### Pattern 19: Iterators vs Lists ```python import sys def process_file_list(filename): """Load entire file into memory.""" with open(filename) as f: lines = f.readlines() # Loads all lines return sum(1 for line in lines if line.strip()) def process_file_iterator(filename): """Process file line by line.""" with open(filename) as f: return sum(1 for line in f if line.strip()) # Iterator uses constant memory # List loads entire file into memory ``` ### Pattern 20: Weakref for Caches ```python import weakref class CachedResource: """Resource that can be garbage collected.""" def __init__(self, data): self.data = data # Regular cache prevents garbage collection regular_cache = {} def get_resource_regular(key): """Get resource from regular cache.""" if key not in regular_cache: regular_cache[key] = CachedResource(f"Data for {key}") return regular_cache[key] # Weak reference cache allows garbage collection weak_cache = weakref.WeakValueDictionary() def get_resource_weak(key): """Get resource from weak cache.""" resource = weak_cache.get(key) if resource is None: resource = CachedResource(f"Data for {key}") weak_cache[key] = resource return resource # When no strong references exist, objects can be GC'd ``` ## Benchmarking Tools ### Custom Benchmark Decorator ```python import time from functools import wraps def benchmark(func): """Decorator to benchmark function execution.""" @wraps(func) def wrapper(*args, **kwargs): start = time.perf_counter() result = func(*args, **kwargs) elapsed = time.perf_counter() - start print(f"{func.__name__} took {elapsed:.6f} seconds") return result return wrapper @benchmark def slow_function(): """Function to benchmark.""" time.sleep(0.5) return sum(range(1000000)) result = slow_function() ``` ### Performance Testing with pytest-benchmark ```python # Install: pip install pytest-benchmark def test_list_comprehension(benchmark): """Benchmark list comprehension.""" result = benchmark(lambda: [i**2 for i in range(10000)]) assert len(result) == 10000 def test_map_function(benchmark): """Benchmark map function.""" result = benchmark(lambda: list(m
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered