Skip to main content
EVOKORE// BROWSE
>

./browse/prompts

10 NODES
🤖system prompt•7 months ago

architecture-decision-records

Write and maintain Architecture Decision Records (ADRs) following

coding
⭐1
# Architecture Decision Records Comprehensive patterns for creating, maintaining, and managing Architecture Decision Records (ADRs) that capture the context and rationale behind significant technical decisions. ## When to Use This Skill - Making significant architectural decisions - Documenting technology choices - Recording design trade-offs - Onboarding new team members - Reviewing historical decisions - Establishing decision-making processes ## Core Concepts ### 1. What is an ADR? An Architecture Decision Record captures: - **Context**: Why we needed to make a decision - **Decision**: What we decided - **Consequences**: What happens as a result ### 2. When to Write an ADR | Write ADR | Skip ADR | | -------------------------- | ---------------------- | | New framework adoption | Minor version upgrades | | Database technology choice | Bug fixes | | API design patterns | Implementation details | | Security architecture | Routine maintenance | | Integration patterns | Configuration changes | ### 3. ADR Lifecycle ``` Proposed → Accepted → Deprecated → Superseded ↓ Rejected ``` ## Templates ### Template 1: Standard ADR (MADR Format) ```markdown # ADR-0001: Use PostgreSQL as Primary Database ## Status Accepted ## Context We need to select a primary database for our new e-commerce platform. The system will handle: - ~10,000 concurrent users - Complex product catalog with hierarchical categories - Transaction processing for orders and payments - Full-text search for products - Geospatial queries for store locator The team has experience with MySQL, PostgreSQL, and MongoDB. We need ACID compliance for financial transactions. ## Decision Drivers - **Must have ACID compliance** for payment processing - **Must support complex queries** for reporting - **Should support full-text search** to reduce infrastructure complexity - **Should have good JSON support** for flexible product attributes - **Team familiarity** reduces onboarding time ## Considered Options ### Option 1: PostgreSQL - **Pros**: ACID compliant, excellent JSON support (JSONB), built-in full-text search, PostGIS for geospatial, team has experience - **Cons**: Slightly more complex replication setup than MySQL ### Option 2: MySQL - **Pros**: Very familiar to team, simple replication, large community - **Cons**: Weaker JSON support, no built-in full-text search (need Elasticsearch), no geospatial without extensions ### Option 3: MongoDB - **Pros**: Flexible schema, native JSON, horizontal scaling - **Cons**: No ACID for multi-document transactions (at decision time), team has limited experience, requires schema design discipline ## Decision We will use **PostgreSQL 15** as our primary database. ## Rationale PostgreSQL provides the best balance of: 1. **ACID compliance** essential for e-commerce transactions 2. **Built-in capabilities** (full-text search, JSONB, PostGIS) reduce infrastructure complexity 3. **Team familiarity** with SQL databases reduces learning curve 4. **Mature ecosystem** with excellent tooling and community support The slight complexity in replication is outweighed by the reduction in additional services (no separate Elasticsearch needed). ## Consequences ### Positive - Single database handles transactions, search, and geospatial queries - Reduced operational complexity (fewer services to manage) - Strong consistency guarantees for financial data - Team can leverage existing SQL expertise ### Negative - Need to learn PostgreSQL-specific features (JSONB, full-text search syntax) - Vertical scaling limits may require read replicas sooner - Some team members need PostgreSQL-specific training ### Risks - Full-text search may not scale as well as dedicated search engines - Mitigation: Design for potential Elasticsearch addition if needed ## Implementation Notes - Use JSONB for flexible product attributes - Implement connection pooling with PgBouncer - Set up streaming replication for read replicas - Use pg_trgm extension for fuzzy search ## Related Decisions - ADR-0002: Caching Strategy (Redis) - complements database choice - ADR-0005: Search Architecture - may supersede if Elasticsearch needed ## References - [PostgreSQL JSON Documentation](https://www.postgresql.org/docs/current/datatype-json.html) - [PostgreSQL Full Text Search](https://www.postgresql.org/docs/current/textsearch.html) - Internal: Performance benchmarks in `/docs/benchmarks/database-comparison.md` ``` ### Template 2: Lightweight ADR ```markdown # ADR-0012: Adopt TypeScript for Frontend Development **Status**: Accepted **Date**: 2024-01-15 **Deciders**: @alice, @bob, @charlie ## Context Our React codebase has grown to 50+ components with increasing bug reports related to prop type mismatches and undefined errors. PropTypes provide runtime-only checking. ## Decision Adopt TypeScript for all new frontend code. Migrate existing code incrementally. ## Consequences **Good**: Catch type errors at compile time, better IDE support, self-documenting code. **Bad**: Learning curve for team, initial slowdown, build complexity increase. **Mitigations**: TypeScript training sessions, allow gradual adoption with `allowJs: true`. ``` ### Template 3: Y-Statement Format ```markdown # ADR-0015: API Gateway Selection In the context of **building a microservices architecture**, facing **the need for centralized API management, authentication, and rate limiting**, we decided for **Kong Gateway** and against **AWS API Gateway and custom Nginx solution**, to achieve **vendor independence, plugin extensibility, and team familiarity with Lua**, accepting that **we need to manage Kong infrastructure ourselves**. ``` ### Template 4: ADR for Deprecation ```markdown # ADR-0020: Deprecate MongoDB in Favor of PostgreSQL ## Status Accepted (Supersedes ADR-0003) ## Context ADR-0003 (2021) chose MongoDB for user profile storage due to schema flexibility needs. Since then: - MongoDB's multi-document transactions remain problematic for our use case - Our schema has stabilized and rarely changes - We now have PostgreSQL expertise from other services - Maintaining two databases increases operational burden ## Decision Deprecate MongoDB and migrate user profiles to PostgreSQL. ## Migration Plan 1. **Phase 1** (Week 1-2): Create PostgreSQL schema, dual-write enabled 2. **Phase 2** (Week 3-4): Backfill historical data, validate consistency 3. **Phase 3** (Week 5): Switch reads to PostgreSQL, monitor 4. **Phase 4** (Week 6): Remove MongoDB writes, decommission ## Consequences ### Positive - Single database technology reduces operational complexity - ACID transactions for user data - Team can focus PostgreSQL expertise ### Negative - Migration effort (~4 weeks) - Risk of data issues during migration - Lose some schema flexibility ## Lessons Learned Document from ADR-0003 experience: - Schema flexibility benefits were overestimated - Operational cost of multiple databases was underestimated - Consider long-term maintenance in technology decisions ``` ### Template 5: Request for Comments (RFC) Style ```markdown # RFC-0025: Adopt Event Sourcing for Order Management ## Summary Propose adopting event sourcing pattern for the order management domain to improve auditability, enable temporal queries, and support business analytics. ## Motivation Current challenges: 1. Audit requirements need complete order history 2. "What was the order state at time X?" queries are impossible 3. Analytics team needs event stream for real-time dashboards 4. Order state reconstruction for customer support is manual ## Detailed Design ### Event Store ``` OrderCreated { orderId, customerId, items[], timestamp } OrderItemAdded { orderId, item, timestamp } OrderItemRemoved { orderId, itemId, timestamp } PaymentReceived { orderId, amount, paymentId, timestamp } OrderShipped { orderId, trackingNumber, timestamp } ``` ### Projections - **CurrentOrderState**: Materialized view for queries - **OrderHistory**: Complete timeline for audit - **DailyOrderMetrics**: Analytics aggregation ### Technology - Event Store: EventStoreDB (purpose-built, handles projections) - Alternative considered: Kafka + custom projection service ## Drawbacks - Learning curve for team - Increased complexity vs. CRUD - Need to design events carefully (immutable once stored) - Storage growth (events never deleted) ## Alternatives 1. **Audit tables**: Simpler but doesn't enable temporal queries 2. **CDC from existing DB**: Complex, doesn't change data model 3. **Hybrid**: Event source only for order state changes ## Unresolved Questions - [ ] Event schema versioning strategy - [ ] Retention policy for events - [ ] Snapshot frequency for performance ## Implementation Plan 1. Prototype with single order type (2 weeks) 2. Team training on event sourcing (1 week) 3. Full implementation and migration (4 weeks) 4. Monitoring and optimization (ongoing) ## References - [Event Sourcing by Martin Fowler](https://martinfowler.com/eaaDev/EventSourcing.html) - [EventStoreDB Documentation](https://www.eventstore.com/docs) ``` ## ADR Management ### Directory Structure ``` docs/ ├── adr/ │ ├── README.md # Index and guidelines │ ├── template.md # Team's ADR template │ ├── 0001-use-postgresql.md │ ├── 0002-caching-strategy.md │ ├── 0003-mongodb-user-profiles.md # [DEPRECATED] │ └── 0020-deprecate-mongodb.md # Supersedes 0003 ``` ### ADR Index (README.md) ```markdown # Architecture Decision Records This directory contains Architecture Decision Records (ADRs) for [Project Name]. ## Index | ADR | Title | Status | Date | | ------------------------------------- | ---------------------------------- | ---------- | ---------- | | [0001](0001-use-postgresql.md) | Use PostgreSQL as Primary Database | Accepted | 2024-01-10 | | [0002](0002-caching-strategy.md) | Caching Strategy with Redis | Accepted | 2024-01-12 | | [0003](0003-mongodb-user-profiles.md) | MongoDB for User Profiles | Deprecated | 2023-06-15 | | [0020](0020-deprecate-mongodb.md) | Deprecate MongoDB | Accepted | 2024-01-15 | ## Creating a New ADR 1. Copy `template.md` to `NNNN-title-with-dashes.md` 2. Fill in the template 3. Submit PR for review 4. Update this index after approval ## ADR Status - **Proposed**: Under discussion - **Accepted**: Decision made, implementing - **Deprecated**: No longer relevant - **Superseded**: Replaced by another ADR - **Rejected**: Considered but not adopted ``` ### Automation (adr-tools) ```bash # Install adr-tools brew install adr-tools # Initialize ADR directory adr init docs/adr # Create new ADR adr new "Use PostgreSQL as Primary Database" # Supersede an ADR adr new -s 3 "Deprecate MongoDB in Favor of PostgreSQL" # Generate table of contents adr generate toc > docs/adr/README.md # Link related ADRs adr link 2 "Complements" 1 "Is complemented by" ``` ## Review Process ```markdown ## ADR Review Checklist ### Before Submission - [ ] Context clearly explains the problem - [ ] All viable options considered - [ ] Pros/cons balanced and honest - [ ] Consequences (positive and negative) documented - [ ] Related ADRs linked ### During Review - [ ] At least 2 senior engineers reviewed - [ ] Affected teams consulted - [ ] Security implications considered - [ ] Cost implications documented - [ ] Reversibility assessed ### After Acceptance - [ ] ADR index updated - [ ] Team notified - [ ] Implementation tickets created - [ ] Related documentation updated ``` ## Best Practices ### Do's - **Write ADRs early** - Before implementation starts - **Keep them short** - 1-2 pages maximum - **Be honest about trade-offs** - Include real cons - **Link related decisions** - Build decision graph - **Update status** - Deprecate when superseded ### Don'ts - **Don't change accepted ADRs** - Write new ones to supersede - **Don't skip context** - Future readers need background - **Don't hide failures** - Rejected decisions are valuable - **Don't be vague** - Specific decisions, specific consequences - **Don't forget implementation** - ADR without action is waste ## Resources - [Documenting Architecture Decisions (Michael Nygard)](https://cognitect.com/blog/2011/11/15/documenting-architecture-decisions) - [MADR Template](https://adr.github.io/madr/) - [ADR GitHub Organization](https://adr.github.io/) - [adr-tools](https://github.com/npryce/adr-tools)
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

employment-contract-templates

Create employment contracts, offer letters, and HR policy documents

coding
⭐1
# Employment Contract Templates Templates and patterns for creating legally sound employment documentation including contracts, offer letters, and HR policies. ## When to Use This Skill - Drafting employment contracts - Creating offer letters - Writing employee handbooks - Developing HR policies - Standardizing employment documentation - Onboarding documentation ## Core Concepts ### 1. Employment Document Types | Document | Purpose | When Used | | ----------------------- | ----------------------- | ------------- | | **Offer Letter** | Initial job offer | Pre-hire | | **Employment Contract** | Formal agreement | Hire | | **Employee Handbook** | Policies & procedures | Onboarding | | **NDA** | Confidentiality | Before access | | **Non-Compete** | Competition restriction | Hire/Exit | ### 2. Key Legal Considerations ``` Employment Relationship: ├── At-Will vs. Contract ├── Employee vs. Contractor ├── Full-Time vs. Part-Time ├── Exempt vs. Non-Exempt └── Jurisdiction-Specific Requirements ``` **DISCLAIMER: These templates are for informational purposes only and do not constitute legal advice. Consult with qualified legal counsel before using any employment documents.** ## Templates ### Template 1: Offer Letter ```markdown # EMPLOYMENT OFFER LETTER [Company Letterhead] Date: [DATE] [Candidate Name] [Address] [City, State ZIP] Dear [Candidate Name], We are pleased to extend an offer of employment for the position of [JOB TITLE] at [COMPANY NAME]. We believe your skills and experience will be valuable additions to our team. ## Position Details **Title:** [Job Title] **Department:** [Department] **Reports To:** [Manager Name/Title] **Location:** [Office Location / Remote] **Start Date:** [Proposed Start Date] **Employment Type:** [Full-Time/Part-Time], [Exempt/Non-Exempt] ## Compensation **Base Salary:** $[AMOUNT] per [year/hour], paid [bi-weekly/semi-monthly/monthly] **Bonus:** [Eligible for annual bonus of up to X% based on company and individual performance / Not applicable] **Equity:** [X shares of stock options vesting over 4 years with 1-year cliff / Not applicable] ## Benefits You will be eligible for our standard benefits package, including: - Health insurance (medical, dental, vision) effective [date] - 401(k) with [X]% company match - [x] days paid time off per year - [x] paid holidays - [Other benefits] Full details will be provided during onboarding. ## Contingencies This offer is contingent upon: - Successful completion of background check - Verification of your right to work in [Country] - Execution of required employment documents including: - Confidentiality Agreement - [Non-Compete Agreement, if applicable] - [IP Assignment Agreement] ## At-Will Employment Please note that employment with [Company Name] is at-will. This means that either you or the Company may terminate the employment relationship at any time, with or without cause or notice. This offer letter does not constitute a contract of employment for any specific period. ## Acceptance To accept this offer, please sign below and return by [DEADLINE DATE]. This offer will expire if not accepted by that date. We are excited about the possibility of you joining our team. If you have any questions, please contact [HR Contact] at [email/phone]. Sincerely, --- [Hiring Manager Name] [Title] [Company Name] --- ## ACCEPTANCE I accept this offer of employment and agree to the terms stated above. Signature: ************\_************ Printed Name: ************\_************ Date: ************\_************ Anticipated Start Date: ************\_************ ``` ### Template 2: Employment Agreement (Contract Position) ```markdown # EMPLOYMENT AGREEMENT This Employment Agreement ("Agreement") is entered into as of [DATE] ("Effective Date") by and between: **Employer:** [COMPANY LEGAL NAME], a [State] [corporation/LLC] with principal offices at [Address] ("Company") **Employee:** [EMPLOYEE NAME], an individual residing at [Address] ("Employee") ## 1. EMPLOYMENT 1.1 **Position.** The Company agrees to employ Employee as [JOB TITLE], reporting to [Manager Title]. Employee accepts such employment subject to the terms of this Agreement. 1.2 **Duties.** Employee shall perform duties consistent with their position, including but not limited to: - [Primary duty 1] - [Primary duty 2] - [Primary duty 3] - Other duties as reasonably assigned 1.3 **Best Efforts.** Employee agrees to devote their full business time, attention, and best efforts to the Company's business during employment. 1.4 **Location.** Employee's primary work location shall be [Location/Remote]. [Travel requirements, if any.] ## 2. TERM 2.1 **Employment Period.** This Agreement shall commence on [START DATE] and continue until terminated as provided herein. 2.2 **At-Will Employment.** [FOR AT-WILL STATES] Notwithstanding anything herein, employment is at-will and may be terminated by either party at any time, with or without cause or notice. [OR FOR FIXED TERM:] 2.2 **Fixed Term.** This Agreement is for a fixed term of [X] months/years, ending on [END DATE], unless terminated earlier as provided herein or extended by mutual written agreement. ## 3. COMPENSATION 3.1 **Base Salary.** Employee shall receive a base salary of $[AMOUNT] per year, payable in accordance with the Company's standard payroll practices, subject to applicable withholdings. 3.2 **Bonus.** Employee may be eligible for an annual discretionary bonus of up to [X]% of base salary, based on [criteria]. Bonus payments are at Company's sole discretion and require active employment at payment date. 3.3 **Equity.** [If applicable] Subject to Board approval and the Company's equity incentive plan, Employee shall be granted [X shares/options] under the terms of a separate Stock Option Agreement. 3.4 **Benefits.** Employee shall be entitled to participate in benefit plans offered to similarly situated employees, subject to plan terms and eligibility requirements. 3.5 **Expenses.** Company shall reimburse Employee for reasonable business expenses incurred in accordance with Company policy. ## 4. CONFIDENTIALITY 4.1 **Confidential Information.** Employee acknowledges access to confidential and proprietary information including: trade secrets, business plans, customer lists, financial data, technical information, and other non-public information ("Confidential Information"). 4.2 **Non-Disclosure.** During and after employment, Employee shall not disclose, use, or permit use of any Confidential Information except as required for their duties or with prior written consent. 4.3 **Return of Materials.** Upon termination, Employee shall immediately return all Company property and Confidential Information in any form. 4.4 **Survival.** Confidentiality obligations survive termination indefinitely for trade secrets and for [3] years for other Confidential Information. ## 5. INTELLECTUAL PROPERTY 5.1 **Work Product.** All inventions, discoveries, works, and developments created by Employee during employment, relating to Company's business, or using Company resources ("Work Product") shall be Company's sole property. 5.2 **Assignment.** Employee hereby assigns to Company all rights in Work Product, including all intellectual property rights. 5.3 **Assistance.** Employee agrees to execute documents and take actions necessary to perfect Company's rights in Work Product. 5.4 **Prior Inventions.** Attached as Exhibit A is a list of any prior inventions that Employee wishes to exclude from this Agreement. ## 6. NON-COMPETITION AND NON-SOLICITATION [NOTE: Enforceability varies by jurisdiction. Consult local counsel.] 6.1 **Non-Competition.** During employment and for [12] months after termination, Employee shall not, directly or indirectly, engage in any business competitive with Company's business within [Geographic Area]. 6.2 **Non-Solicitation of Customers.** During employment and for [12] months after termination, Employee shall not solicit any customer of the Company for competing products or services. 6.3 **Non-Solicitation of Employees.** During employment and for [12] months after termination, Employee shall not recruit or solicit any Company employee to leave Company employment. ## 7. TERMINATION 7.1 **By Company for Cause.** Company may terminate immediately for Cause, defined as: (a) Material breach of this Agreement (b) Conviction of a felony (c) Fraud, dishonesty, or gross misconduct (d) Failure to perform duties after written notice and cure period 7.2 **By Company Without Cause.** Company may terminate without Cause upon [30] days written notice. 7.3 **By Employee.** Employee may terminate upon [30] days written notice. 7.4 **Severance.** [If applicable] Upon termination without Cause, Employee shall receive [X] weeks base salary as severance, contingent upon execution of a release agreement. 7.5 **Effect of Termination.** Upon termination: - All compensation earned through termination date shall be paid - Unvested equity shall be forfeited - Benefits terminate per plan terms - Sections 4, 5, 6, 8, and 9 survive termination ## 8. GENERAL PROVISIONS 8.1 **Entire Agreement.** This Agreement constitutes the entire agreement and supersedes all prior negotiations, representations, and agreements. 8.2 **Amendments.** This Agreement may be amended only by written agreement signed by both parties. 8.3 **Governing Law.** This Agreement shall be governed by the laws of [State], without regard to conflicts of law principles. 8.4 **Dispute Resolution.** [Arbitration clause or jurisdiction selection] 8.5 **Severability.** If any provision is unenforceable, it shall be modified to the minimum extent necessary, and remaining provisions shall remain in effect. 8.6 **Notices.** Notices shall be in writing and delivered to addresses above. 8.7 **Assignment.** Employee may not assign this Agreement. Company may assign to a successor. 8.8 **Waiver.** Failure to enforce any provision shall not constitute waiver. ## 9. ACKNOWLEDGMENTS Employee acknowledges: - Having read and understood this Agreement - Having opportunity to consult with counsel - Agreeing to all terms voluntarily --- IN WITNESS WHEREOF, the parties have executed this Agreement as of the Effective Date. **[COMPANY NAME]** By: ************\_************ Name: [Authorized Signatory] Title: [Title] Date: ************\_************ **EMPLOYEE** Signature: ************\_************ Name: [Employee Name] Date: ************\_************ --- ## EXHIBIT A: PRIOR INVENTIONS [Employee to list any prior inventions, if any, or write "None"] --- ``` ### Template 3: Employee Handbook Policy Section ```markdown # EMPLOYEE HANDBOOK - POLICY SECTION ## EMPLOYMENT POLICIES ### Equal Employment Opportunity [Company Name] is an equal opportunity employer. We do not discriminate based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. This policy applies to all employment practices including: - Recruitment and hiring - Compensation and benefits - Training and development - Promotions and transfers - Termination ### Anti-Harassment Policy [Company Name] is committed to providing a workplace free from harassment. Harassment based on any protected characteristic is strictly prohibited. **Prohibited Conduct Includes:** - Unwelcome sexual advances or requests for sexual favors - Offensive comments, jokes, or slurs - Physical conduct such as assault or unwanted touching - Visual conduct such as displaying offensive images - Threatening, intimidating, or hostile acts **Reporting Procedure:** 1. Report to your manager, HR, or any member of leadership 2. Reports may be made verbally or in writing 3. Anonymous reports are accepted via [hotline/email] **Investigation:** All reports will be promptly investigated. Retaliation against anyone who reports harassment is strictly prohibited and will result in disciplinary action up to termination. ### Work Hours and Attendance **Standard Hours:** [8:00 AM - 5:00 PM, Monday through Friday] **Core Hours:** [10:00 AM - 3:00 PM] - Employees expected to be available **Flexible Work:** [Policy on remote work, flexible scheduling] **Attendance Expectations:** - Notify your manager as soon as possible if you will be absent - Excessive unexcused absences may result in disciplinary action - [x] unexcused absences in [Y] days considered excessive ### Paid Time Off (PTO) **PTO Accrual:** | Years of Service | Annual PTO Days | |------------------|-----------------| | 0-2 years | 15 days | | 3-5 years | 20 days | | 6+ years | 25 days | **PTO Guidelines:** - PTO accrues per pay period - Maximum accrual: [X] days (use it or lose it after) - Request PTO at least [2] weeks in advance - Manager approval required - PTO may not be taken during [blackout periods] ### Sick Leave - [x] days sick leave per year - May be used for personal illness or family member care - Doctor's note required for absences exceeding [3] days ### Holidays The following paid holidays are observed: - New Year's Day - Martin Luther King Jr. Day - Presidents Day - Memorial Day - Independence Day - Labor Day - Thanksgiving Day - Day after Thanksgiving - Christmas Day - [Floating holiday] ### Code of Conduct All employees are expected to: - Act with integrity and honesty - Treat colleagues, customers, and partners with respect - Protect company confidential information - Avoid conflicts of interest - Comply with all laws and regulations - Report any violations of this code **Violations may result in disciplinary action up to and including termination.** ### Technology and Communication **Acceptable Use:** - Company technology is for business purposes - Limited personal use is permitted if it doesn't interfere with work - No illegal activities or viewing inappropriate content **Monitoring:** - Company reserves the right to monitor company systems - Employees should have no expectation of privacy on company devices **Security:** - Use strong passwords and enable 2FA - Report security incidents immediately - Lock devices when unattended ### Social Media Policy **Personal Social Media:** - Clearly state opinions are your own, not the company's - Do not share confidential company information - Be respectful and professional **Company Social Media:** - Only authorized personnel may post on behalf of the company - Follow brand guidelines - Escalate negative comments to [Marketing/PR] --- ## ACKNOWLEDGMENT I acknowledge that I have received a copy of the Employee Handbook and understand that: 1. I am responsible for reading and understanding its contents 2. The handbook does not create a contract of employment 3. Policies may be changed at any time at the company's discretion 4. Employment is at-will [if applicable] I agree to abide by the policies and procedures outlined in this handbook. Employee Signature: ************\_************ Employee Name (Print): ************\_************ Date: ************\_************ ``` ## Best Practices ### Do's - **Consult legal counsel** - Employment law varies by jurisdiction - **Keep copies signed** - Document all agreements - **Update regularly** - Laws and policies change - **Be clear and specific** - Avoid ambiguity - **Train managers** - On policies and procedures ### Don'ts - **Don't use generic templates** - Customize for your jurisdiction - **Don't make promises** - That could create implied contracts - **Don't discriminate** - In language or application - **Don't forget at-will language** - Where applicable - **Don't skip review** - Have legal counsel review all documents ## Resources - [SHRM Employment Templates](https://www.shrm.org/) - [Department of Labor](https://www.dol.gov/) - [EEOC Guidance](https://www.eeoc.gov/) - State-specific labor departments
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

gdpr-data-handling

Implement GDPR-compliant data handling with consent management,

data
⭐1
# GDPR Data Handling Practical implementation guide for GDPR-compliant data processing, consent management, and privacy controls. ## When to Use This Skill - Building systems that process EU personal data - Implementing consent management - Handling data subject requests (DSRs) - Conducting GDPR compliance reviews - Designing privacy-first architectures - Creating data processing agreements ## Core Concepts ### 1. Personal Data Categories | Category | Examples | Protection Level | | ---------------------- | --------------------------- | ------------------ | | **Basic** | Name, email, phone | Standard | | **Sensitive (Art. 9)** | Health, religion, ethnicity | Explicit consent | | **Criminal (Art. 10)** | Convictions, offenses | Official authority | | **Children's** | Under 16 data | Parental consent | ### 2. Legal Bases for Processing ``` Article 6 - Lawful Bases: ├── Consent: Freely given, specific, informed ├── Contract: Necessary for contract performance ├── Legal Obligation: Required by law ├── Vital Interests: Protecting someone's life ├── Public Interest: Official functions └── Legitimate Interest: Balanced against rights ``` ### 3. Data Subject Rights ``` Right to Access (Art. 15) ─┐ Right to Rectification (Art. 16) │ Right to Erasure (Art. 17) │ Must respond Right to Restrict (Art. 18) │ within 1 month Right to Portability (Art. 20) │ Right to Object (Art. 21) ─┘ ``` ## Implementation Patterns ### Pattern 1: Consent Management ```javascript // Consent data model const consentSchema = { userId: String, consents: [ { purpose: String, // 'marketing', 'analytics', etc. granted: Boolean, timestamp: Date, source: String, // 'web_form', 'api', etc. version: String, // Privacy policy version ipAddress: String, // For proof userAgent: String, // For proof }, ], auditLog: [ { action: String, // 'granted', 'withdrawn', 'updated' purpose: String, timestamp: Date, source: String, }, ], }; // Consent service class ConsentManager { async recordConsent(userId, purpose, granted, metadata) { const consent = { purpose, granted, timestamp: new Date(), source: metadata.source, version: await this.getCurrentPolicyVersion(), ipAddress: metadata.ipAddress, userAgent: metadata.userAgent, }; // Store consent await this.db.consents.updateOne( { userId }, { $push: { consents: consent, auditLog: { action: granted ? "granted" : "withdrawn", purpose, timestamp: consent.timestamp, source: metadata.source, }, }, }, { upsert: true }, ); // Emit event for downstream systems await this.eventBus.emit("consent.changed", { userId, purpose, granted, timestamp: consent.timestamp, }); } async hasConsent(userId, purpose) { const record = await this.db.consents.findOne({ userId }); if (!record) return false; const latestConsent = record.consents .filter((c) => c.purpose === purpose) .sort((a, b) => b.timestamp - a.timestamp)[0]; return latestConsent?.granted === true; } async getConsentHistory(userId) { const record = await this.db.consents.findOne({ userId }); return record?.auditLog || []; } } ``` ```html <!-- GDPR-compliant consent UI --> <div class="consent-banner" role="dialog" aria-labelledby="consent-title"> <h2 id="consent-title">Cookie Preferences</h2> <p> We use cookies to improve your experience. Select your preferences below. </p> <form id="consent-form"> <!-- Necessary - always on, no consent needed --> <div class="consent-category"> <input type="checkbox" id="necessary" checked disabled /> <label for="necessary"> <strong>Necessary</strong> <span>Required for the website to function. Cannot be disabled.</span> </label> </div> <!-- Analytics - requires consent --> <div class="consent-category"> <input type="checkbox" id="analytics" name="analytics" /> <label for="analytics"> <strong>Analytics</strong> <span>Help us understand how you use our site.</span> </label> </div> <!-- Marketing - requires consent --> <div class="consent-category"> <input type="checkbox" id="marketing" name="marketing" /> <label for="marketing"> <strong>Marketing</strong> <span>Personalized ads based on your interests.</span> </label> </div> <div class="consent-actions"> <button type="button" id="accept-all">Accept All</button> <button type="button" id="reject-all">Reject All</button> <button type="submit">Save Preferences</button> </div> <p class="consent-links"> <a href="/privacy-policy">Privacy Policy</a> | <a href="/cookie-policy">Cookie Policy</a> </p> </form> </div> ``` ### Pattern 2: Data Subject Access Request (DSAR) ```python from datetime import datetime, timedelta from typing import Dict, List, Optional import json class DSARHandler: """Handle Data Subject Access Requests.""" RESPONSE_DEADLINE_DAYS = 30 EXTENSION_ALLOWED_DAYS = 60 # For complex requests def __init__(self, data_sources: List['DataSource']): self.data_sources = data_sources async def submit_request( self, request_type: str, # 'access', 'erasure', 'rectification', 'portability' user_id: str, verified: bool, details: Optional[Dict] = None ) -> str: """Submit a new DSAR.""" request = { 'id': self.generate_request_id(), 'type': request_type, 'user_id': user_id, 'status': 'pending_verification' if not verified else 'processing', 'submitted_at': datetime.utcnow(), 'deadline': datetime.utcnow() + timedelta(days=self.RESPONSE_DEADLINE_DAYS), 'details': details or {}, 'audit_log': [{ 'action': 'submitted', 'timestamp': datetime.utcnow(), 'details': 'Request received' }] } await self.db.dsar_requests.insert_one(request) await self.notify_dpo(request) return request['id'] async def process_access_request(self, request_id: str) -> Dict: """Process a data access request.""" request = await self.get_request(request_id) if request['type'] != 'access': raise ValueError("Not an access request") # Collect data from all sources user_data = {} for source in self.data_sources: try: data = await source.get_user_data(request['user_id']) user_data[source.name] = data except Exception as e: user_data[source.name] = {'error': str(e)} # Format response response = { 'request_id': request_id, 'generated_at': datetime.utcnow().isoformat(), 'data_categories': list(user_data.keys()), 'data': user_data, 'retention_info': await self.get_retention_info(), 'processing_purposes': await self.get_processing_purposes(), 'third_party_recipients': await self.get_recipients() } # Update request status await self.update_request(request_id, 'completed', response) return response async def process_erasure_request(self, request_id: str) -> Dict: """Process a right to erasure request.""" request = await self.get_request(request_id) if request['type'] != 'erasure': raise ValueError("Not an erasure request") results = {} exceptions = [] for source in self.data_sources: try: # Check for legal exceptions can_delete, reason = await source.can_delete(request['user_id']) if can_delete: await source.delete_user_data(request['user_id']) results[source.name] = 'deleted' else: exceptions.append({ 'source': source.name, 'reason': reason # e.g., 'legal retention requirement' }) results[source.name] = f'retained: {reason}' except Exception as e: results[source.name] = f'error: {str(e)}' response = { 'request_id': request_id, 'completed_at': datetime.utcnow().isoformat(), 'results': results, 'exceptions': exceptions } await self.update_request(request_id, 'completed', response) return response async def process_portability_request(self, request_id: str) -> bytes: """Generate portable data export.""" request = await self.get_request(request_id) user_data = await self.process_access_request(request_id) # Convert to machine-readable format (JSON) portable_data = { 'export_date': datetime.utcnow().isoformat(), 'format_version': '1.0', 'data': user_data['data'] } return json.dumps(portable_data, indent=2, default=str).encode() ``` ### Pattern 3: Data Retention ```python from datetime import datetime, timedelta from enum import Enum class RetentionBasis(Enum): CONSENT = "consent" CONTRACT = "contract" LEGAL_OBLIGATION = "legal_obligation" LEGITIMATE_INTEREST = "legitimate_interest" class DataRetentionPolicy: """Define and enforce data retention policies.""" POLICIES = { 'user_account': { 'retention_period_days': 365 * 3, # 3 years after last activity 'basis': RetentionBasis.CONTRACT, 'trigger': 'last_activity_date', 'archive_before_delete': True }, 'transaction_records': { 'retention_period_days': 365 * 7, # 7 years for tax 'basis': RetentionBasis.LEGAL_OBLIGATION, 'trigger': 'transaction_date', 'archive_before_delete': True, 'legal_reference': 'Tax regulations require 7 year retention' }, 'marketing_consent': { 'retention_period_days': 365 * 2, # 2 years 'basis': RetentionBasis.CONSENT, 'trigger': 'consent_date', 'archive_before_delete': False }, 'support_tickets': { 'retention_period_days': 365 * 2, 'basis': RetentionBasis.LEGITIMATE_INTEREST, 'trigger': 'ticket_closed_date', 'archive_before_delete': True }, 'analytics_data': { 'retention_period_days': 365, # 1 year 'basis': RetentionBasis.CONSENT, 'trigger': 'collection_date', 'archive_before_delete': False, 'anonymize_instead': True } } async def apply_retention_policies(self): """Run retention policy enforcement.""" for data_type, policy in self.POLICIES.items(): cutoff_date = datetime.utcnow() - timedelta( days=policy['retention_period_days'] ) if policy.get('anonymize_instead'): await self.anonymize_old_data(data_type, cutoff_date) else: if policy.get('archive_before_delete'): await self.archive_data(data_type, cutoff_date) await self.delete_old_data(data_type, cutoff_date) await self.log_retention_action(data_type, cutoff_date) async def anonymize_old_data(self, data_type: str, before_date: datetime): """Anonymize data instead of deleting.""" # Example: Replace identifying fields with hashes if data_type == 'analytics_data': await self.db.analytics.update_many( {'collection_date': {'$lt': before_date}}, {'$set': { 'user_id': None, 'ip_address': None, 'device_id': None, 'anonymized': True, 'anonymized_date': datetime.utcnow() }} ) ``` ### Pattern 4: Privacy by Design ```python class PrivacyFirstDataModel: """Example of privacy-by-design data model.""" # Separate PII from behavioral data user_profile_schema = { 'user_id': str, # UUID, not sequential 'email_hash': str, # Hashed for lookups 'created_at': datetime, # Minimal data collection 'preferences': { 'language': str, 'timezone': str } } # Encrypted at rest user_pii_schema = { 'user_id': str, 'email': str, # Encrypted 'name': str, # Encrypted 'phone': str, # Encrypted (optional) 'address': dict, # Encrypted (optional) 'encryption_key_id': str } # Pseudonymized behavioral data analytics_schema = { 'session_id': str, # Not linked to user_id 'pseudonym_id': str, # Rotating pseudonym 'events': list, 'device_category': str, # Generalized, not specific 'country': str, # Not city-level } class DataMinimization: """Implement data minimization principles.""" @staticmethod def collect_only_needed(form_data: dict, purpose: str) -> dict: """Filter form data to only fields needed for purpose.""" REQUIRED_FIELDS = { 'account_creation': ['email', 'password'], 'newsletter': ['email'], 'purchase': ['email', 'name', 'address', 'payment'], 'support': ['email', 'message'] } allowed = REQUIRED_FIELDS.get(purpose, []) return {k: v for k, v in form_data.items() if k in allowed} @staticmethod def generalize_location(ip_address: str) -> str: """Generalize IP to country level only.""" import geoip2.database reader = geoip2.database.Reader('GeoLite2-Country.mmdb') try: response = reader.country(ip_address) return response.country.iso_code except: return 'UNKNOWN' ``` ### Pattern 5: Breach Notification ```python from datetime import datetime from enum import Enum class BreachSeverity(Enum): LOW = "low" MEDIUM = "medium" HIGH = "high" CRITICAL = "critical" class BreachNotificationHandler: """Handle GDPR breach notification requirements.""" AUTHORITY_NOTIFICATION_HOURS = 72 AFFECTED_NOTIFICATION_REQUIRED_SEVERITY = BreachSeverity.HIGH async def report_breach( self, description: str, data_types: List[str], affected_count: int, severity: BreachSeverity ) -> dict: """Report and handle a data breach.""" breach = { 'id': self.generate_breach_id(), 'reported_at': datetime.utcnow(), 'description': description, 'data_types_affected': data_types, 'affected_individuals_count': affected_count, 'severity': severity.value, 'status': 'investigating', 'timeline': [{ 'event': 'breach_reported', 'timestamp': datetime.utcnow(), 'details': description }] } await self.db.breaches.insert_one(breach) # Immediate notifications await self.notify_dpo(breach) await self.notify_security_team(breach) # Authority notification required within 72 hours if self.requires_authority_notification(severity, data_types): breach['authority_notification_deadline'] = ( datetime.utcnow() + timedelta(hours=self.AUTHORITY_NOTIFICATION_HOURS) ) await self.schedule_authority_notification(breach) # Affected individuals notification if severity.value in [BreachSeverity.HIGH.value, BreachSeverity.CRITICAL.value]: await self.schedule_individual_notifications(breach) return breach def requires_authority_notification( self, severity: BreachSeverity, data_types: List[str] ) -> bool: """Determine if supervisory authority must be notified.""" # Always notify for sensitive data sensitive_types = ['health', 'financial', 'credentials', 'biometric'] if any(t in sensitive_types for t in data_types): return True # Notify for medium+ severity return severity in [BreachSeverity.MEDIUM, BreachSeverity.HIGH, BreachSeverity.CRITICAL] async def generate_authority_report(self, breach_id: str) -> dict: """Generate report for supervisory authority.""" breach = await self.get_breach(breach_id) return { 'organization': { 'name': self.config.org_name, 'contact': self.config.dpo_contact, 'registration': self.config.registration_number }, 'breach': { 'nature': breach['description'], 'categories_affected': breach['data_types_affected'], 'approximate_number_affected': breach['affected_individuals_count'], 'likely_consequences': self.assess_consequences(breach), 'measures_taken': await self.get_remediation_measures(breach_id), 'measures_proposed': await self.get_proposed_measures(breach_id) }, 'timeline': breach['timeline'], 'submitted_at': datetime.utcnow().isoformat() } ``` ## Compliance Checklist ```markdown ## GDPR Implementation Checklist ### Legal Basis - [ ] Documented legal basis for each processing activity - [ ] Consent mechanisms meet GDPR requirements - [ ] Legitimate interest assessments completed ### Transparency - [ ] Privacy policy is clear and accessible - [ ] Processing purposes clearly stated - [ ] Data retention periods documented ### Data Subject Rights - [ ] Access request process implemented - [ ] Erasure request process implemented - [ ] Portability export available - [ ] Rectification process available - [ ] Response within 30-day deadline ### Security - [ ] Encryption at rest implemented - [ ] Encryption in transit (TLS) - [ ] Access controls in place - [ ] Audit logging enabled ### Breach Response - [ ] Breach detection mechanisms - [ ] 72-hour notification process - [ ] Breach documentation system ### Documentation - [ ] Records of processing activities (Art. 30) - [ ] Data protection impact assessments - [ ] Data processing agreements with vendors ``` ## Best Practices ### Do's - **Minimize data collection** - Only collect what's needed - **Document everything** - Processing activities, legal bases - **Encrypt PII** - At rest and in transit - **Implement access controls** - Need-to-know basis - **Regular audits** - Verify compliance continuously ### Don'ts - **Don't pre-check consent boxes** - Must be opt-in - **Don't bundle consent** - Separate purposes separately - **Don't retain indefinitely** - Defi
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

incident-runbook-templates

Create structured incident response runbooks with step-by-step

coding
⭐1
# Incident Runbook Templates Production-ready templates for incident response runbooks covering detection, triage, mitigation, resolution, and communication. ## When to Use This Skill - Creating incident response procedures - Building service-specific runbooks - Establishing escalation paths - Documenting recovery procedures - Responding to active incidents - Onboarding on-call engineers ## Core Concepts ### 1. Incident Severity Levels | Severity | Impact | Response Time | Example | | -------- | -------------------------- | ----------------- | ----------------------- | | **SEV1** | Complete outage, data loss | 15 min | Production down | | **SEV2** | Major degradation | 30 min | Critical feature broken | | **SEV3** | Minor impact | 2 hours | Non-critical bug | | **SEV4** | Minimal impact | Next business day | Cosmetic issue | ### 2. Runbook Structure ``` 1. Overview & Impact 2. Detection & Alerts 3. Initial Triage 4. Mitigation Steps 5. Root Cause Investigation 6. Resolution Procedures 7. Verification & Rollback 8. Communication Templates 9. Escalation Matrix ``` ## Runbook Templates ### Template 1: Service Outage Runbook ````markdown # [Service Name] Outage Runbook ## Overview **Service**: Payment Processing Service **Owner**: Platform Team **Slack**: #payments-incidents **PagerDuty**: payments-oncall ## Impact Assessment - [ ] Which customers are affected? - [ ] What percentage of traffic is impacted? - [ ] Are there financial implications? - [ ] What's the blast radius? ## Detection ### Alerts - `payment_error_rate > 5%` (PagerDuty) - `payment_latency_p99 > 2s` (Slack) - `payment_success_rate < 95%` (PagerDuty) ### Dashboards - [Payment Service Dashboard](https://grafana/d/payments) - [Error Tracking](https://sentry.io/payments) - [Dependency Status](https://status.stripe.com) ## Initial Triage (First 5 Minutes) ### 1. Assess Scope ```bash # Check service health kubectl get pods -n payments -l app=payment-service # Check recent deployments kubectl rollout history deployment/payment-service -n payments # Check error rates curl -s "http://prometheus:9090/api/v1/query?query=sum(rate(http_requests_total{status=~'5..'}[5m]))" ``` ```` ### 2. Quick Health Checks - [ ] Can you reach the service? `curl -I https://api.company.com/payments/health` - [ ] Database connectivity? Check connection pool metrics - [ ] External dependencies? Check Stripe, bank API status - [ ] Recent changes? Check deploy history ### 3. Initial Classification | Symptom | Likely Cause | Go To Section | | -------------------- | ------------------- | ------------- | | All requests failing | Service down | Section 4.1 | | High latency | Database/dependency | Section 4.2 | | Partial failures | Code bug | Section 4.3 | | Spike in errors | Traffic surge | Section 4.4 | ## Mitigation Procedures ### 4.1 Service Completely Down ```bash # Step 1: Check pod status kubectl get pods -n payments # Step 2: If pods are crash-looping, check logs kubectl logs -n payments -l app=payment-service --tail=100 # Step 3: Check recent deployments kubectl rollout history deployment/payment-service -n payments # Step 4: ROLLBACK if recent deploy is suspect kubectl rollout undo deployment/payment-service -n payments # Step 5: Scale up if resource constrained kubectl scale deployment/payment-service -n payments --replicas=10 # Step 6: Verify recovery kubectl rollout status deployment/payment-service -n payments ``` ### 4.2 High Latency ```bash # Step 1: Check database connections kubectl exec -n payments deploy/payment-service -- \ curl localhost:8080/metrics | grep db_pool # Step 2: Check slow queries (if DB issue) psql -h $DB_HOST -U $DB_USER -c " SELECT pid, now() - query_start AS duration, query FROM pg_stat_activity WHERE state = 'active' AND duration > interval '5 seconds' ORDER BY duration DESC;" # Step 3: Kill long-running queries if needed psql -h $DB_HOST -U $DB_USER -c "SELECT pg_terminate_backend(pid);" # Step 4: Check external dependency latency curl -w "@curl-format.txt" -o /dev/null -s https://api.stripe.com/v1/health # Step 5: Enable circuit breaker if dependency is slow kubectl set env deployment/payment-service \ STRIPE_CIRCUIT_BREAKER_ENABLED=true -n payments ``` ### 4.3 Partial Failures (Specific Errors) ```bash # Step 1: Identify error pattern kubectl logs -n payments -l app=payment-service --tail=500 | \ grep -i error | sort | uniq -c | sort -rn | head -20 # Step 2: Check error tracking # Go to Sentry: https://sentry.io/payments # Step 3: If specific endpoint, enable feature flag to disable curl -X POST https://api.company.com/internal/feature-flags \ -d '{"flag": "DISABLE_PROBLEMATIC_FEATURE", "enabled": true}' # Step 4: If data issue, check recent data changes psql -h $DB_HOST -c " SELECT * FROM audit_log WHERE table_name = 'payment_methods' AND created_at > now() - interval '1 hour';" ``` ### 4.4 Traffic Surge ```bash # Step 1: Check current request rate kubectl top pods -n payments # Step 2: Scale horizontally kubectl scale deployment/payment-service -n payments --replicas=20 # Step 3: Enable rate limiting kubectl set env deployment/payment-service \ RATE_LIMIT_ENABLED=true \ RATE_LIMIT_RPS=1000 -n payments # Step 4: If attack, block suspicious IPs kubectl apply -f - <<EOF apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: block-suspicious namespace: payments spec: podSelector: matchLabels: app: payment-service ingress: - from: - ipBlock: cidr: 0.0.0.0/0 except: - 192.168.1.0/24 # Suspicious range EOF ``` ## Verification Steps ```bash # Verify service is healthy curl -s https://api.company.com/payments/health | jq # Verify error rate is back to normal curl -s "http://prometheus:9090/api/v1/query?query=sum(rate(http_requests_total{status=~'5..'}[5m]))" | jq '.data.result[0].value[1]' # Verify latency is acceptable curl -s "http://prometheus:9090/api/v1/query?query=histogram_quantile(0.99,sum(rate(http_request_duration_seconds_bucket[5m]))by(le))" | jq # Smoke test critical flows ./scripts/smoke-test-payments.sh ``` ## Rollback Procedures ```bash # Rollback Kubernetes deployment kubectl rollout undo deployment/payment-service -n payments # Rollback database migration (if applicable) ./scripts/db-rollback.sh $MIGRATION_VERSION # Rollback feature flag curl -X POST https://api.company.com/internal/feature-flags \ -d '{"flag": "NEW_PAYMENT_FLOW", "enabled": false}' ``` ## Escalation Matrix | Condition | Escalate To | Contact | | ----------------------------- | ------------------- | ------------------- | | > 15 min unresolved SEV1 | Engineering Manager | @manager (Slack) | | Data breach suspected | Security Team | #security-incidents | | Financial impact > $10k | Finance + Legal | @finance-oncall | | Customer communication needed | Support Lead | @support-lead | ## Communication Templates ### Initial Notification (Internal) ``` 🚨 INCIDENT: Payment Service Degradation Severity: SEV2 Status: Investigating Impact: ~20% of payment requests failing Start Time: [TIME] Incident Commander: [NAME] Current Actions: - Investigating root cause - Scaling up service - Monitoring dashboards Updates in #payments-incidents ``` ### Status Update ``` 📊 UPDATE: Payment Service Incident Status: Mitigating Impact: Reduced to ~5% failure rate Duration: 25 minutes Actions Taken: - Rolled back deployment v2.3.4 → v2.3.3 - Scaled service from 5 → 10 replicas Next Steps: - Continuing to monitor - Root cause analysis in progress ETA to Resolution: ~15 minutes ``` ### Resolution Notification ``` ✅ RESOLVED: Payment Service Incident Duration: 45 minutes Impact: ~5,000 affected transactions Root Cause: Memory leak in v2.3.4 Resolution: - Rolled back to v2.3.3 - Transactions auto-retried successfully Follow-up: - Postmortem scheduled for [DATE] - Bug fix in progress ``` ```` ### Template 2: Database Incident Runbook ```markdown # Database Incident Runbook ## Quick Reference | Issue | Command | |-------|---------| | Check connections | `SELECT count(*) FROM pg_stat_activity;` | | Kill query | `SELECT pg_terminate_backend(pid);` | | Check replication lag | `SELECT extract(epoch from (now() - pg_last_xact_replay_timestamp()));` | | Check locks | `SELECT * FROM pg_locks WHERE NOT granted;` | ## Connection Pool Exhaustion ```sql -- Check current connections SELECT datname, usename, state, count(*) FROM pg_stat_activity GROUP BY datname, usename, state ORDER BY count(*) DESC; -- Identify long-running connections SELECT pid, usename, datname, state, query_start, query FROM pg_stat_activity WHERE state != 'idle' ORDER BY query_start; -- Terminate idle connections SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE state = 'idle' AND query_start < now() - interval '10 minutes'; ```` ## Replication Lag ```sql -- Check lag on replica SELECT CASE WHEN pg_last_wal_receive_lsn() = pg_last_wal_replay_lsn() THEN 0 ELSE extract(epoch from now() - pg_last_xact_replay_timestamp()) END AS lag_seconds; -- If lag > 60s, consider: -- 1. Check network between primary/replica -- 2. Check replica disk I/O -- 3. Consider failover if unrecoverable ``` ## Disk Space Critical ```bash # Check disk usage df -h /var/lib/postgresql/data # Find large tables psql -c "SELECT relname, pg_size_pretty(pg_total_relation_size(relid)) FROM pg_catalog.pg_statio_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 10;" # VACUUM to reclaim space psql -c "VACUUM FULL large_table;" # If emergency, delete old data or expand disk ``` ``` ## Best Practices ### Do's - **Keep runbooks updated** - Review after every incident - **Test runbooks regularly** - Game days, chaos engineering - **Include rollback steps** - Always have an escape hatch - **Document assumptions** - What must be true for steps to work - **Link to dashboards** - Quick access during stress ### Don'ts - **Don't assume knowledge** - Write for 3 AM brain - **Don't skip verification** - Confirm each step worked - **Don't forget communication** - Keep stakeholders informed - **Don't work alone** - Escalate early - **Don't skip postmortems** - Learn from every incident ## Resources - [Google SRE Book - Incident Management](https://sre.google/sre-book/managing-incidents/) - [PagerDuty Incident Response](https://response.pagerduty.com/) - [Atlassian Incident Management](https://www.atlassian.com/incident-management) ```
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

embedding-strategies

Select and optimize embedding models for semantic search and RAG

coding
⭐1
# Embedding Strategies Guide to selecting and optimizing embedding models for vector search applications. ## When to Use This Skill - Choosing embedding models for RAG - Optimizing chunking strategies - Fine-tuning embeddings for domains - Comparing embedding model performance - Reducing embedding dimensions - Handling multilingual content ## Core Concepts ### 1. Embedding Model Comparison (2026) | Model | Dimensions | Max Tokens | Best For | | -------------------------- | ---------- | ---------- | ----------------------------------- | | **voyage-3-large** | 1024 | 32000 | Claude apps (Anthropic recommended) | | **voyage-3** | 1024 | 32000 | Claude apps, cost-effective | | **voyage-code-3** | 1024 | 32000 | Code search | | **voyage-finance-2** | 1024 | 32000 | Financial documents | | **voyage-law-2** | 1024 | 32000 | Legal documents | | **text-embedding-3-large** | 3072 | 8191 | OpenAI apps, high accuracy | | **text-embedding-3-small** | 1536 | 8191 | OpenAI apps, cost-effective | | **bge-large-en-v1.5** | 1024 | 512 | Open source, local deployment | | **all-MiniLM-L6-v2** | 384 | 256 | Fast, lightweight | | **multilingual-e5-large** | 1024 | 512 | Multi-language | ### 2. Embedding Pipeline ``` Document → Chunking → Preprocessing → Embedding Model → Vector ↓ [Overlap, Size] [Clean, Normalize] [API/Local] ``` ## Templates ### Template 1: Voyage AI Embeddings (Recommended for Claude) ```python from langchain_voyageai import VoyageAIEmbeddings from typing import List import os # Initialize Voyage AI embeddings (recommended by Anthropic for Claude) embeddings = VoyageAIEmbeddings( model="voyage-3-large", voyage_api_key=os.environ.get("VOYAGE_API_KEY") ) def get_embeddings(texts: List[str]) -> List[List[float]]: """Get embeddings from Voyage AI.""" return embeddings.embed_documents(texts) def get_query_embedding(query: str) -> List[float]: """Get single query embedding.""" return embeddings.embed_query(query) # Specialized models for domains code_embeddings = VoyageAIEmbeddings(model="voyage-code-3") finance_embeddings = VoyageAIEmbeddings(model="voyage-finance-2") legal_embeddings = VoyageAIEmbeddings(model="voyage-law-2") ``` ### Template 2: OpenAI Embeddings ```python from openai import OpenAI from typing import List import numpy as np client = OpenAI() def get_embeddings( texts: List[str], model: str = "text-embedding-3-small", dimensions: int = None ) -> List[List[float]]: """Get embeddings from OpenAI with optional dimension reduction.""" # Handle batching for large lists batch_size = 100 all_embeddings = [] for i in range(0, len(texts), batch_size): batch = texts[i:i + batch_size] kwargs = {"input": batch, "model": model} if dimensions: # Matryoshka dimensionality reduction kwargs["dimensions"] = dimensions response = client.embeddings.create(**kwargs) embeddings = [item.embedding for item in response.data] all_embeddings.extend(embeddings) return all_embeddings def get_embedding(text: str, **kwargs) -> List[float]: """Get single embedding.""" return get_embeddings([text], **kwargs)[0] # Dimension reduction with Matryoshka embeddings def get_reduced_embedding(text: str, dimensions: int = 512) -> List[float]: """Get embedding with reduced dimensions (Matryoshka).""" return get_embedding( text, model="text-embedding-3-small", dimensions=dimensions ) ``` ### Template 3: Local Embeddings with Sentence Transformers ```python from sentence_transformers import SentenceTransformer from typing import List, Optional import numpy as np class LocalEmbedder: """Local embedding with sentence-transformers.""" def __init__( self, model_name: str = "BAAI/bge-large-en-v1.5", device: str = "cuda" ): self.model = SentenceTransformer(model_name, device=device) self.model_name = model_name def embed( self, texts: List[str], normalize: bool = True, show_progress: bool = False ) -> np.ndarray: """Embed texts with optional normalization.""" embeddings = self.model.encode( texts, normalize_embeddings=normalize, show_progress_bar=show_progress, convert_to_numpy=True ) return embeddings def embed_query(self, query: str) -> np.ndarray: """Embed a query with appropriate prefix for retrieval models.""" # BGE and similar models benefit from query prefix if "bge" in self.model_name.lower(): query = f"Represent this sentence for searching relevant passages: {query}" return self.embed([query])[0] def embed_documents(self, documents: List[str]) -> np.ndarray: """Embed documents for indexing.""" return self.embed(documents) # E5 model with instructions class E5Embedder: def __init__(self, model_name: str = "intfloat/multilingual-e5-large"): self.model = SentenceTransformer(model_name) def embed_query(self, query: str) -> np.ndarray: """E5 requires 'query:' prefix for queries.""" return self.model.encode(f"query: {query}") def embed_document(self, document: str) -> np.ndarray: """E5 requires 'passage:' prefix for documents.""" return self.model.encode(f"passage: {document}") ``` ### Template 4: Chunking Strategies ```python from typing import List, Tuple import re def chunk_by_tokens( text: str, chunk_size: int = 512, chunk_overlap: int = 50, tokenizer=None ) -> List[str]: """Chunk text by token count.""" import tiktoken tokenizer = tokenizer or tiktoken.get_encoding("cl100k_base") tokens = tokenizer.encode(text) chunks = [] start = 0 while start < len(tokens): end = start + chunk_size chunk_tokens = tokens[start:end] chunk_text = tokenizer.decode(chunk_tokens) chunks.append(chunk_text) start = end - chunk_overlap return chunks def chunk_by_sentences( text: str, max_chunk_size: int = 1000, min_chunk_size: int = 100 ) -> List[str]: """Chunk text by sentences, respecting size limits.""" import nltk sentences = nltk.sent_tokenize(text) chunks = [] current_chunk = [] current_size = 0 for sentence in sentences: sentence_size = len(sentence) if current_size + sentence_size > max_chunk_size and current_chunk: chunks.append(" ".join(current_chunk)) current_chunk = [] current_size = 0 current_chunk.append(sentence) current_size += sentence_size if current_chunk: chunks.append(" ".join(current_chunk)) return chunks def chunk_by_semantic_sections( text: str, headers_pattern: str = r'^#{1,3}\s+.+$' ) -> List[Tuple[str, str]]: """Chunk markdown by headers, preserving hierarchy.""" lines = text.split('\n') chunks = [] current_header = "" current_content = [] for line in lines: if re.match(headers_pattern, line, re.MULTILINE): if current_content: chunks.append((current_header, '\n'.join(current_content))) current_header = line current_content = [] else: current_content.append(line) if current_content: chunks.append((current_header, '\n'.join(current_content))) return chunks def recursive_character_splitter( text: str, chunk_size: int = 1000, chunk_overlap: int = 200, separators: List[str] = None ) -> List[str]: """LangChain-style recursive splitter.""" separators = separators or ["\n\n", "\n", ". ", " ", ""] def split_text(text: str, separators: List[str]) -> List[str]: if not text: return [] separator = separators[0] remaining_separators = separators[1:] if separator == "": # Character-level split return [text[i:i+chunk_size] for i in range(0, len(text), chunk_size - chunk_overlap)] splits = text.split(separator) chunks = [] current_chunk = [] current_length = 0 for split in splits: split_length = len(split) + len(separator) if current_length + split_length > chunk_size and current_chunk: chunk_text = separator.join(current_chunk) # Recursively split if still too large if len(chunk_text) > chunk_size and remaining_separators: chunks.extend(split_text(chunk_text, remaining_separators)) else: chunks.append(chunk_text) # Start new chunk with overlap overlap_splits = [] overlap_length = 0 for s in reversed(current_chunk): if overlap_length + len(s) <= chunk_overlap: overlap_splits.insert(0, s) overlap_length += len(s) else: break current_chunk = overlap_splits current_length = overlap_length current_chunk.append(split) current_length += split_length if current_chunk: chunks.append(separator.join(current_chunk)) return chunks return split_text(text, separators) ``` ### Template 5: Domain-Specific Embedding Pipeline ```python import re from typing import List, Optional from dataclasses import dataclass @dataclass class EmbeddedDocument: id: str document_id: str chunk_index: int text: str embedding: List[float] metadata: dict class DomainEmbeddingPipeline: """Pipeline for domain-specific embeddings.""" def __init__( self, embedding_model: str = "voyage-3-large", chunk_size: int = 512, chunk_overlap: int = 50, preprocessing_fn=None ): self.embeddings = VoyageAIEmbeddings(model=embedding_model) self.chunk_size = chunk_size self.chunk_overlap = chunk_overlap self.preprocess = preprocessing_fn or self._default_preprocess def _default_preprocess(self, text: str) -> str: """Default preprocessing.""" # Remove excessive whitespace text = re.sub(r'\s+', ' ', text) # Remove special characters (customize for your domain) text = re.sub(r'[^\w\s.,!?-]', '', text) return text.strip() async def process_documents( self, documents: List[dict], id_field: str = "id", content_field: str = "content", metadata_fields: Optional[List[str]] = None ) -> List[EmbeddedDocument]: """Process documents for vector storage.""" processed = [] for doc in documents: content = doc[content_field] doc_id = doc[id_field] # Preprocess cleaned = self.preprocess(content) # Chunk chunks = chunk_by_tokens( cleaned, self.chunk_size, self.chunk_overlap ) # Create embeddings embeddings = await self.embeddings.aembed_documents(chunks) # Create records for i, (chunk, embedding) in enumerate(zip(chunks, embeddings)): metadata = {"document_id": doc_id, "chunk_index": i} # Add specified metadata fields if metadata_fields: for field in metadata_fields: if field in doc: metadata[field] = doc[field] processed.append(EmbeddedDocument( id=f"{doc_id}_chunk_{i}", document_id=doc_id, chunk_index=i, text=chunk, embedding=embedding, metadata=metadata )) return processed # Code-specific pipeline class CodeEmbeddingPipeline: """Specialized pipeline for code embeddings.""" def __init__(self): # Use Voyage's code-specific model self.embeddings = VoyageAIEmbeddings(model="voyage-code-3") def chunk_code(self, code: str, language: str) -> List[dict]: """Chunk code by functions/classes using tree-sitter.""" try: import tree_sitter_languages parser = tree_sitter_languages.get_parser(language) tree = parser.parse(bytes(code, "utf8")) chunks = [] # Extract function and class definitions self._extract_nodes(tree.root_node, code, chunks) return chunks except ImportError: # Fallback to simple chunking return [{"text": code, "type": "module"}] def _extract_nodes(self, node, source_code: str, chunks: list): """Recursively extract function/class definitions.""" if node.type in ['function_definition', 'class_definition', 'method_definition']: text = source_code[node.start_byte:node.end_byte] chunks.append({ "text": text, "type": node.type, "name": self._get_name(node), "start_line": node.start_point[0], "end_line": node.end_point[0] }) for child in node.children: self._extract_nodes(child, source_code, chunks) def _get_name(self, node) -> str: """Extract name from function/class node.""" for child in node.children: if child.type == 'identifier' or child.type == 'name': return child.text.decode('utf8') return "unknown" async def embed_with_context( self, chunk: str, context: str = "" ) -> List[float]: """Embed code with surrounding context.""" if context: combined = f"Context: {context}\n\nCode:\n{chunk}" else: combined = chunk return await self.embeddings.aembed_query(combined) ``` ### Template 6: Embedding Quality Evaluation ```python import numpy as np from typing import List, Dict def evaluate_retrieval_quality( queries: List[str], relevant_docs: List[List[str]], # List of relevant doc IDs per query retrieved_docs: List[List[str]], # List of retrieved doc IDs per query k: int = 10 ) -> Dict[str, float]: """Evaluate embedding quality for retrieval.""" def precision_at_k(relevant: set, retrieved: List[str], k: int) -> float: retrieved_k = retrieved[:k] relevant_retrieved = len(set(retrieved_k) & relevant) return relevant_retrieved / k if k > 0 else 0 def recall_at_k(relevant: set, retrieved: List[str], k: int) -> float: retrieved_k = retrieved[:k] relevant_retrieved = len(set(retrieved_k) & relevant) return relevant_retrieved / len(relevant) if relevant else 0 def mrr(relevant: set, retrieved: List[str]) -> float: for i, doc in enumerate(retrieved): if doc in relevant: return 1 / (i + 1) return 0 def ndcg_at_k(relevant: set, retrieved: List[str], k: int) -> float: dcg = sum( 1 / np.log2(i + 2) if doc in relevant else 0 for i, doc in enumerate(retrieved[:k]) ) ideal_dcg = sum(1 / np.log2(i + 2) for i in range(min(len(relevant), k))) return dcg / ideal_dcg if ideal_dcg > 0 else 0 metrics = { f"precision@{k}": [], f"recall@{k}": [], "mrr": [], f"ndcg@{k}": [] } for relevant, retrieved in zip(relevant_docs, retrieved_docs): relevant_set = set(relevant) metrics[f"precision@{k}"].append(precision_at_k(relevant_set, retrieved, k)) metrics[f"recall@{k}"].append(recall_at_k(relevant_set, retrieved, k)) metrics["mrr"].append(mrr(relevant_set, retrieved)) metrics[f"ndcg@{k}"].append(ndcg_at_k(relevant_set, retrieved, k)) return {name: np.mean(values) for name, values in metrics.items()} def compute_embedding_similarity( embeddings1: np.ndarray, embeddings2: np.ndarray, metric: str = "cosine" ) -> np.ndarray: """Compute similarity matrix between embedding sets.""" if metric == "cosine": # Normalize and compute dot product norm1 = embeddings1 / np.linalg.norm(embeddings1, axis=1, keepdims=True) norm2 = embeddings2 / np.linalg.norm(embeddings2, axis=1, keepdims=True) return norm1 @ norm2.T elif metric == "euclidean": from scipy.spatial.distance import cdist return -cdist(embeddings1, embeddings2, metric='euclidean') elif metric == "dot": return embeddings1 @ embeddings2.T else: raise ValueError(f"Unknown metric: {metric}") def compare_embedding_models( texts: List[str], models: Dict[str, callable], queries: List[str], relevant_indices: List[List[int]], k: int = 5 ) -> Dict[str, Dict[str, float]]: """Compare multiple embedding models on retrieval quality.""" results = {} for model_name, embed_fn in models.items(): # Embed all texts doc_embeddings = np.array(embed_fn(texts)) retrieved_per_query = [] for query in queries: query_embedding = np.array(embed_fn([query])[0]) # Compute similarities similarities = compute_embedding_similarity( query_embedding.reshape(1, -1), doc_embeddings, metric="cosine" )[0] # Get top-k indices top_k_indices = np.argsort(similarities)[::-1][:k] retrieved_per_query.append([str(i) for i in top_k_indices]) # Convert relevant indices to string IDs relevant_docs = [[str(i) for i in indices] for indices in relevant_indices] results[model_name] = evaluate_retrieval_quality( queries, relevant_docs, retrieved_per_query, k ) return results ``` ## Best Practices ### Do's - **Match model to use case**: Code vs prose vs multilingual - **Chunk thoughtfully**: Preserve semantic boundaries - **Normalize embeddings**: For cosine similarity search - **Batch requests**: More efficient than one-by-one - **Cache embeddings**: Avoid recomputing for static content - **Use Voyage AI for Claude apps**: Recommended by Anthropic ### Don'ts - **Don't ignore token limits**: Truncation loses information - **Don't mix embedding models**: Incompatible vector spaces - **Don't skip preprocessing**: Garbage in, garbage out - **Don't over-chunk**: Lose important context - **Don't forget metadata**: Essential for filtering and debugging ## Resources - [Voyage AI Documentation](https://docs.voyageai.com/) - [OpenAI Embeddings Guide](https://platform.openai.com/docs/guides/embeddings) - [Sentence Transformers](https://www.sbert.net/) - [MTEB Benchmar
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

prompt-engineering-patterns

Master advanced prompt engineering techniques to maximize LLM

coding
⭐1
# Prompt Engineering Patterns Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability. ## When to Use This Skill - Designing complex prompts for production LLM applications - Optimizing prompt performance and consistency - Implementing structured reasoning patterns (chain-of-thought, tree-of-thought) - Building few-shot learning systems with dynamic example selection - Creating reusable prompt templates with variable interpolation - Debugging and refining prompts that produce inconsistent outputs - Implementing system prompts for specialized AI assistants - Using structured outputs (JSON mode) for reliable parsing ## Core Capabilities ### 1. Few-Shot Learning - Example selection strategies (semantic similarity, diversity sampling) - Balancing example count with context window constraints - Constructing effective demonstrations with input-output pairs - Dynamic example retrieval from knowledge bases - Handling edge cases through strategic example selection ### 2. Chain-of-Thought Prompting - Step-by-step reasoning elicitation - Zero-shot CoT with "Let's think step by step" - Few-shot CoT with reasoning traces - Self-consistency techniques (sampling multiple reasoning paths) - Verification and validation steps ### 3. Structured Outputs - JSON mode for reliable parsing - Pydantic schema enforcement - Type-safe response handling - Error handling for malformed outputs ### 4. Prompt Optimization - Iterative refinement workflows - A/B testing prompt variations - Measuring prompt performance metrics (accuracy, consistency, latency) - Reducing token usage while maintaining quality - Handling edge cases and failure modes ### 5. Template Systems - Variable interpolation and formatting - Conditional prompt sections - Multi-turn conversation templates - Role-based prompt composition - Modular prompt components ### 6. System Prompt Design - Setting model behavior and constraints - Defining output formats and structure - Establishing role and expertise - Safety guidelines and content policies - Context setting and background information ## Quick Start ```python from langchain_anthropic import ChatAnthropic from langchain_core.prompts import ChatPromptTemplate from pydantic import BaseModel, Field # Define structured output schema class SQLQuery(BaseModel): query: str = Field(description="The SQL query") explanation: str = Field(description="Brief explanation of what the query does") tables_used: list[str] = Field(description="List of tables referenced") # Initialize model with structured output llm = ChatAnthropic(model="claude-sonnet-4-6") structured_llm = llm.with_structured_output(SQLQuery) # Create prompt template prompt = ChatPromptTemplate.from_messages([ ("system", """You are an expert SQL developer. Generate efficient, secure SQL queries. Always use parameterized queries to prevent SQL injection. Explain your reasoning briefly."""), ("user", "Convert this to SQL: {query}") ]) # Create chain chain = prompt | structured_llm # Use result = await chain.ainvoke({ "query": "Find all users who registered in the last 30 days" }) print(result.query) print(result.explanation) ``` ## Key Patterns ### Pattern 1: Structured Output with Pydantic ```python from anthropic import Anthropic from pydantic import BaseModel, Field from typing import Literal import json class SentimentAnalysis(BaseModel): sentiment: Literal["positive", "negative", "neutral"] confidence: float = Field(ge=0, le=1) key_phrases: list[str] reasoning: str async def analyze_sentiment(text: str) -> SentimentAnalysis: """Analyze sentiment with structured output.""" client = Anthropic() message = client.messages.create( model="claude-sonnet-4-6", max_tokens=500, messages=[{ "role": "user", "content": f"""Analyze the sentiment of this text. Text: {text} Respond with JSON matching this schema: {{ "sentiment": "positive" | "negative" | "neutral", "confidence": 0.0-1.0, "key_phrases": ["phrase1", "phrase2"], "reasoning": "brief explanation" }}""" }] ) return SentimentAnalysis(**json.loads(message.content[0].text)) ``` ### Pattern 2: Chain-of-Thought with Self-Verification ```python from langchain_core.prompts import ChatPromptTemplate cot_prompt = ChatPromptTemplate.from_template(""" Solve this problem step by step. Problem: {problem} Instructions: 1. Break down the problem into clear steps 2. Work through each step showing your reasoning 3. State your final answer 4. Verify your answer by checking it against the original problem Format your response as: ## Steps [Your step-by-step reasoning] ## Answer [Your final answer] ## Verification [Check that your answer is correct] """) ``` ### Pattern 3: Few-Shot with Dynamic Example Selection ```python from langchain_voyageai import VoyageAIEmbeddings from langchain_core.example_selectors import SemanticSimilarityExampleSelector from langchain_chroma import Chroma # Create example selector with semantic similarity example_selector = SemanticSimilarityExampleSelector.from_examples( examples=[ {"input": "How do I reset my password?", "output": "Go to Settings > Security > Reset Password"}, {"input": "Where can I see my order history?", "output": "Navigate to Account > Orders"}, {"input": "How do I contact support?", "output": "Click Help > Contact Us or email support@example.com"}, ], embeddings=VoyageAIEmbeddings(model="voyage-3-large"), vectorstore_cls=Chroma, k=2 # Select 2 most similar examples ) async def get_few_shot_prompt(query: str) -> str: """Build prompt with dynamically selected examples.""" examples = await example_selector.aselect_examples({"input": query}) examples_text = "\n".join( f"User: {ex['input']}\nAssistant: {ex['output']}" for ex in examples ) return f"""You are a helpful customer support assistant. Here are some example interactions: {examples_text} Now respond to this query: User: {query} Assistant:""" ``` ### Pattern 4: Progressive Disclosure Start with simple prompts, add complexity only when needed: ```python PROMPT_LEVELS = { # Level 1: Direct instruction "simple": "Summarize this article: {text}", # Level 2: Add constraints "constrained": """Summarize this article in 3 bullet points, focusing on: - Key findings - Main conclusions - Practical implications Article: {text}""", # Level 3: Add reasoning "reasoning": """Read this article carefully. 1. First, identify the main topic and thesis 2. Then, extract the key supporting points 3. Finally, summarize in 3 bullet points Article: {text} Summary:""", # Level 4: Add examples "few_shot": """Read articles and provide concise summaries. Example: Article: "New research shows that regular exercise can reduce anxiety by up to 40%..." Summary: • Regular exercise reduces anxiety by up to 40% • 30 minutes of moderate activity 3x/week is sufficient • Benefits appear within 2 weeks of starting Now summarize this article: Article: {text} Summary:""" } ``` ### Pattern 5: Error Recovery and Fallback ```python from pydantic import BaseModel, ValidationError import json class ResponseWithConfidence(BaseModel): answer: str confidence: float sources: list[str] alternative_interpretations: list[str] = [] ERROR_RECOVERY_PROMPT = """ Answer the question based on the context provided. Context: {context} Question: {question} Instructions: 1. If you can answer confidently (>0.8), provide a direct answer 2. If you're somewhat confident (0.5-0.8), provide your best answer with caveats 3. If you're uncertain (<0.5), explain what information is missing 4. Always provide alternative interpretations if the question is ambiguous Respond in JSON: {{ "answer": "your answer or 'I cannot determine this from the context'", "confidence": 0.0-1.0, "sources": ["relevant context excerpts"], "alternative_interpretations": ["if question is ambiguous"] }} """ async def answer_with_fallback( context: str, question: str, llm ) -> ResponseWithConfidence: """Answer with error recovery and fallback.""" prompt = ERROR_RECOVERY_PROMPT.format(context=context, question=question) try: response = await llm.ainvoke(prompt) return ResponseWithConfidence(**json.loads(response.content)) except (json.JSONDecodeError, ValidationError) as e: # Fallback: try to extract answer without structure simple_prompt = f"Based on: {context}\n\nAnswer: {question}" simple_response = await llm.ainvoke(simple_prompt) return ResponseWithConfidence( answer=simple_response.content, confidence=0.5, sources=["fallback extraction"], alternative_interpretations=[] ) ``` ### Pattern 6: Role-Based System Prompts ```python SYSTEM_PROMPTS = { "analyst": """You are a senior data analyst with expertise in SQL, Python, and business intelligence. Your responsibilities: - Write efficient, well-documented queries - Explain your analysis methodology - Highlight key insights and recommendations - Flag any data quality concerns Communication style: - Be precise and technical when discussing methodology - Translate technical findings into business impact - Use clear visualizations when helpful""", "assistant": """You are a helpful AI assistant focused on accuracy and clarity. Core principles: - Always cite sources when making factual claims - Acknowledge uncertainty rather than guessing - Ask clarifying questions when the request is ambiguous - Provide step-by-step explanations for complex topics Constraints: - Do not provide medical, legal, or financial advice - Redirect harmful requests appropriately - Protect user privacy""", "code_reviewer": """You are a senior software engineer conducting code reviews. Review criteria: - Correctness: Does the code work as intended? - Security: Are there any vulnerabilities? - Performance: Are there efficiency concerns? - Maintainability: Is the code readable and well-structured? - Best practices: Does it follow language idioms? Output format: 1. Summary assessment (approve/request changes) 2. Critical issues (must fix) 3. Suggestions (nice to have) 4. Positive feedback (what's done well)""" } ``` ## Integration Patterns ### With RAG Systems ```python RAG_PROMPT = """You are a knowledgeable assistant that answers questions based on provided context. Context (retrieved from knowledge base): {context} Instructions: 1. Answer ONLY based on the provided context 2. If the context doesn't contain the answer, say "I don't have information about that in my knowledge base" 3. Cite specific passages using [1], [2] notation 4. If the question is ambiguous, ask for clarification Question: {question} Answer:""" ``` ### With Validation and Verification ```python VALIDATED_PROMPT = """Complete the following task: Task: {task} After generating your response, verify it meets ALL these criteria: ✓ Directly addresses the original request ✓ Contains no factual errors ✓ Is appropriately detailed (not too brief, not too verbose) ✓ Uses proper formatting ✓ Is safe and appropriate If verification fails on any criterion, revise before responding. Response:""" ``` ## Performance Optimization ### Token Efficiency ```python # Before: Verbose prompt (150+ tokens) verbose_prompt = """ I would like you to please take the following text and provide me with a comprehensive summary of the main points. The summary should capture the key ideas and important details while being concise and easy to understand. """ # After: Concise prompt (30 tokens) concise_prompt = """Summarize the key points concisely: {text} Summary:""" ``` ### Caching Common Prefixes ```python from anthropic import Anthropic client = Anthropic() # Use prompt caching for repeated system prompts response = client.messages.create( model="claude-sonnet-4-6", max_tokens=1000, system=[ { "type": "text", "text": LONG_SYSTEM_PROMPT, "cache_control": {"type": "ephemeral"} } ], messages=[{"role": "user", "content": user_query}] ) ``` ## Best Practices 1. **Be Specific**: Vague prompts produce inconsistent results 2. **Show, Don't Tell**: Examples are more effective than descriptions 3. **Use Structured Outputs**: Enforce schemas with Pydantic for reliability 4. **Test Extensively**: Evaluate on diverse, representative inputs 5. **Iterate Rapidly**: Small changes can have large impacts 6. **Monitor Performance**: Track metrics in production 7. **Version Control**: Treat prompts as code with proper versioning 8. **Document Intent**: Explain why prompts are structured as they are ## Common Pitfalls - **Over-engineering**: Starting with complex prompts before trying simple ones - **Example pollution**: Using examples that don't match the target task - **Context overflow**: Exceeding token limits with excessive examples - **Ambiguous instructions**: Leaving room for multiple interpretations - **Ignoring edge cases**: Not testing on unusual or boundary inputs - **No error handling**: Assuming outputs will always be well-formed - **Hardcoded values**: Not parameterizing prompts for reuse ## Success Metrics Track these KPIs for your prompts: - **Accuracy**: Correctness of outputs - **Consistency**: Reproducibility across similar inputs - **Latency**: Response time (P50, P95, P99) - **Token Usage**: Average tokens per request - **Success Rate**: Percentage of valid, parseable outputs - **User Satisfaction**: Ratings and feedback ## Resources - [Anthropic Prompt Engineering Guide](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering) - [Claude Prompt Caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) - [OpenAI Prompt Engineering](https://platform.openai.com/docs/guides/prompt-engineering) - [LangChain Prompts](https://python.langchain.com/docs/concepts/prompts/)
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

risk-metrics-calculation

Calculate portfolio risk metrics including VaR, CVaR, Sharpe,

coding
⭐1
# Risk Metrics Calculation Comprehensive risk measurement toolkit for portfolio management, including Value at Risk, Expected Shortfall, and drawdown analysis. ## When to Use This Skill - Measuring portfolio risk - Implementing risk limits - Building risk dashboards - Calculating risk-adjusted returns - Setting position sizes - Regulatory reporting ## Core Concepts ### 1. Risk Metric Categories | Category | Metrics | Use Case | | ----------------- | --------------- | -------------------- | | **Volatility** | Std Dev, Beta | General risk | | **Tail Risk** | VaR, CVaR | Extreme losses | | **Drawdown** | Max DD, Calmar | Capital preservation | | **Risk-Adjusted** | Sharpe, Sortino | Performance | ### 2. Time Horizons ``` Intraday: Minute/hourly VaR for day traders Daily: Standard risk reporting Weekly: Rebalancing decisions Monthly: Performance attribution Annual: Strategic allocation ``` ## Implementation ### Pattern 1: Core Risk Metrics ```python import numpy as np import pandas as pd from scipy import stats from typing import Dict, Optional, Tuple class RiskMetrics: """Core risk metric calculations.""" def __init__(self, returns: pd.Series, rf_rate: float = 0.02): """ Args: returns: Series of periodic returns rf_rate: Annual risk-free rate """ self.returns = returns self.rf_rate = rf_rate self.ann_factor = 252 # Trading days per year # Volatility Metrics def volatility(self, annualized: bool = True) -> float: """Standard deviation of returns.""" vol = self.returns.std() if annualized: vol *= np.sqrt(self.ann_factor) return vol def downside_deviation(self, threshold: float = 0, annualized: bool = True) -> float: """Standard deviation of returns below threshold.""" downside = self.returns[self.returns < threshold] if len(downside) == 0: return 0.0 dd = downside.std() if annualized: dd *= np.sqrt(self.ann_factor) return dd def beta(self, market_returns: pd.Series) -> float: """Beta relative to market.""" aligned = pd.concat([self.returns, market_returns], axis=1).dropna() if len(aligned) < 2: return np.nan cov = np.cov(aligned.iloc[:, 0], aligned.iloc[:, 1]) return cov[0, 1] / cov[1, 1] if cov[1, 1] != 0 else 0 # Value at Risk def var_historical(self, confidence: float = 0.95) -> float: """Historical VaR at confidence level.""" return -np.percentile(self.returns, (1 - confidence) * 100) def var_parametric(self, confidence: float = 0.95) -> float: """Parametric VaR assuming normal distribution.""" z_score = stats.norm.ppf(confidence) return self.returns.mean() - z_score * self.returns.std() def var_cornish_fisher(self, confidence: float = 0.95) -> float: """VaR with Cornish-Fisher expansion for non-normality.""" z = stats.norm.ppf(confidence) s = stats.skew(self.returns) # Skewness k = stats.kurtosis(self.returns) # Excess kurtosis # Cornish-Fisher expansion z_cf = (z + (z**2 - 1) * s / 6 + (z**3 - 3*z) * k / 24 - (2*z**3 - 5*z) * s**2 / 36) return -(self.returns.mean() + z_cf * self.returns.std()) # Conditional VaR (Expected Shortfall) def cvar(self, confidence: float = 0.95) -> float: """Expected Shortfall / CVaR / Average VaR.""" var = self.var_historical(confidence) return -self.returns[self.returns <= -var].mean() # Drawdown Analysis def drawdowns(self) -> pd.Series: """Calculate drawdown series.""" cumulative = (1 + self.returns).cumprod() running_max = cumulative.cummax() return (cumulative - running_max) / running_max def max_drawdown(self) -> float: """Maximum drawdown.""" return self.drawdowns().min() def avg_drawdown(self) -> float: """Average drawdown.""" dd = self.drawdowns() return dd[dd < 0].mean() if (dd < 0).any() else 0 def drawdown_duration(self) -> Dict[str, int]: """Drawdown duration statistics.""" dd = self.drawdowns() in_drawdown = dd < 0 # Find drawdown periods drawdown_starts = in_drawdown & ~in_drawdown.shift(1).fillna(False) drawdown_ends = ~in_drawdown & in_drawdown.shift(1).fillna(False) durations = [] current_duration = 0 for i in range(len(dd)): if in_drawdown.iloc[i]: current_duration += 1 elif current_duration > 0: durations.append(current_duration) current_duration = 0 if current_duration > 0: durations.append(current_duration) return { "max_duration": max(durations) if durations else 0, "avg_duration": np.mean(durations) if durations else 0, "current_duration": current_duration } # Risk-Adjusted Returns def sharpe_ratio(self) -> float: """Annualized Sharpe ratio.""" excess_return = self.returns.mean() * self.ann_factor - self.rf_rate vol = self.volatility(annualized=True) return excess_return / vol if vol > 0 else 0 def sortino_ratio(self) -> float: """Sortino ratio using downside deviation.""" excess_return = self.returns.mean() * self.ann_factor - self.rf_rate dd = self.downside_deviation(threshold=0, annualized=True) return excess_return / dd if dd > 0 else 0 def calmar_ratio(self) -> float: """Calmar ratio (return / max drawdown).""" annual_return = (1 + self.returns).prod() ** (self.ann_factor / len(self.returns)) - 1 max_dd = abs(self.max_drawdown()) return annual_return / max_dd if max_dd > 0 else 0 def omega_ratio(self, threshold: float = 0) -> float: """Omega ratio.""" returns_above = self.returns[self.returns > threshold] - threshold returns_below = threshold - self.returns[self.returns <= threshold] if returns_below.sum() == 0: return np.inf return returns_above.sum() / returns_below.sum() # Information Ratio def information_ratio(self, benchmark_returns: pd.Series) -> float: """Information ratio vs benchmark.""" active_returns = self.returns - benchmark_returns tracking_error = active_returns.std() * np.sqrt(self.ann_factor) active_return = active_returns.mean() * self.ann_factor return active_return / tracking_error if tracking_error > 0 else 0 # Summary def summary(self) -> Dict[str, float]: """Generate comprehensive risk summary.""" dd_stats = self.drawdown_duration() return { # Returns "total_return": (1 + self.returns).prod() - 1, "annual_return": (1 + self.returns).prod() ** (self.ann_factor / len(self.returns)) - 1, # Volatility "annual_volatility": self.volatility(), "downside_deviation": self.downside_deviation(), # VaR & CVaR "var_95_historical": self.var_historical(0.95), "var_99_historical": self.var_historical(0.99), "cvar_95": self.cvar(0.95), # Drawdowns "max_drawdown": self.max_drawdown(), "avg_drawdown": self.avg_drawdown(), "max_drawdown_duration": dd_stats["max_duration"], # Risk-Adjusted "sharpe_ratio": self.sharpe_ratio(), "sortino_ratio": self.sortino_ratio(), "calmar_ratio": self.calmar_ratio(), "omega_ratio": self.omega_ratio(), # Distribution "skewness": stats.skew(self.returns), "kurtosis": stats.kurtosis(self.returns), } ``` ### Pattern 2: Portfolio Risk ```python class PortfolioRisk: """Portfolio-level risk calculations.""" def __init__( self, returns: pd.DataFrame, weights: Optional[pd.Series] = None ): """ Args: returns: DataFrame with asset returns (columns = assets) weights: Portfolio weights (default: equal weight) """ self.returns = returns self.weights = weights if weights is not None else \ pd.Series(1/len(returns.columns), index=returns.columns) self.ann_factor = 252 def portfolio_return(self) -> float: """Weighted portfolio return.""" return (self.returns @ self.weights).mean() * self.ann_factor def portfolio_volatility(self) -> float: """Portfolio volatility.""" cov_matrix = self.returns.cov() * self.ann_factor port_var = self.weights @ cov_matrix @ self.weights return np.sqrt(port_var) def marginal_risk_contribution(self) -> pd.Series: """Marginal contribution to risk by asset.""" cov_matrix = self.returns.cov() * self.ann_factor port_vol = self.portfolio_volatility() # Marginal contribution mrc = (cov_matrix @ self.weights) / port_vol return mrc def component_risk(self) -> pd.Series: """Component contribution to total risk.""" mrc = self.marginal_risk_contribution() return self.weights * mrc def risk_parity_weights(self, target_vol: float = None) -> pd.Series: """Calculate risk parity weights.""" from scipy.optimize import minimize n = len(self.returns.columns) cov_matrix = self.returns.cov() * self.ann_factor def risk_budget_objective(weights): port_vol = np.sqrt(weights @ cov_matrix @ weights) mrc = (cov_matrix @ weights) / port_vol rc = weights * mrc target_rc = port_vol / n # Equal risk contribution return np.sum((rc - target_rc) ** 2) constraints = [ {"type": "eq", "fun": lambda w: np.sum(w) - 1}, # Weights sum to 1 ] bounds = [(0.01, 1.0) for _ in range(n)] # Min 1%, max 100% x0 = np.array([1/n] * n) result = minimize( risk_budget_objective, x0, method="SLSQP", bounds=bounds, constraints=constraints ) return pd.Series(result.x, index=self.returns.columns) def correlation_matrix(self) -> pd.DataFrame: """Asset correlation matrix.""" return self.returns.corr() def diversification_ratio(self) -> float: """Diversification ratio (higher = more diversified).""" asset_vols = self.returns.std() * np.sqrt(self.ann_factor) weighted_vol = (self.weights * asset_vols).sum() port_vol = self.portfolio_volatility() return weighted_vol / port_vol if port_vol > 0 else 1 def tracking_error(self, benchmark_returns: pd.Series) -> float: """Tracking error vs benchmark.""" port_returns = self.returns @ self.weights active_returns = port_returns - benchmark_returns return active_returns.std() * np.sqrt(self.ann_factor) def conditional_correlation( self, threshold_percentile: float = 10 ) -> pd.DataFrame: """Correlation during stress periods.""" port_returns = self.returns @ self.weights threshold = np.percentile(port_returns, threshold_percentile) stress_mask = port_returns <= threshold return self.returns[stress_mask].corr() ``` ### Pattern 3: Rolling Risk Metrics ```python class RollingRiskMetrics: """Rolling window risk calculations.""" def __init__(self, returns: pd.Series, window: int = 63): """ Args: returns: Return series window: Rolling window size (default: 63 = ~3 months) """ self.returns = returns self.window = window def rolling_volatility(self, annualized: bool = True) -> pd.Series: """Rolling volatility.""" vol = self.returns.rolling(self.window).std() if annualized: vol *= np.sqrt(252) return vol def rolling_sharpe(self, rf_rate: float = 0.02) -> pd.Series: """Rolling Sharpe ratio.""" rolling_return = self.returns.rolling(self.window).mean() * 252 rolling_vol = self.rolling_volatility() return (rolling_return - rf_rate) / rolling_vol def rolling_var(self, confidence: float = 0.95) -> pd.Series: """Rolling historical VaR.""" return self.returns.rolling(self.window).apply( lambda x: -np.percentile(x, (1 - confidence) * 100), raw=True ) def rolling_max_drawdown(self) -> pd.Series: """Rolling maximum drawdown.""" def max_dd(returns): cumulative = (1 + returns).cumprod() running_max = cumulative.cummax() drawdowns = (cumulative - running_max) / running_max return drawdowns.min() return self.returns.rolling(self.window).apply(max_dd, raw=False) def rolling_beta(self, market_returns: pd.Series) -> pd.Series: """Rolling beta vs market.""" def calc_beta(window_data): port_ret = window_data.iloc[:, 0] mkt_ret = window_data.iloc[:, 1] cov = np.cov(port_ret, mkt_ret) return cov[0, 1] / cov[1, 1] if cov[1, 1] != 0 else 0 combined = pd.concat([self.returns, market_returns], axis=1) return combined.rolling(self.window).apply( lambda x: calc_beta(x.to_frame()), raw=False ).iloc[:, 0] def volatility_regime( self, low_threshold: float = 0.10, high_threshold: float = 0.20 ) -> pd.Series: """Classify volatility regime.""" vol = self.rolling_volatility() def classify(v): if v < low_threshold: return "low" elif v > high_threshold: return "high" else: return "normal" return vol.apply(classify) ``` ### Pattern 4: Stress Testing ```python class StressTester: """Historical and hypothetical stress testing.""" # Historical crisis periods HISTORICAL_SCENARIOS = { "2008_financial_crisis": ("2008-09-01", "2009-03-31"), "2020_covid_crash": ("2020-02-19", "2020-03-23"), "2022_rate_hikes": ("2022-01-01", "2022-10-31"), "dot_com_bust": ("2000-03-01", "2002-10-01"), "flash_crash_2010": ("2010-05-06", "2010-05-06"), } def __init__(self, returns: pd.Series, weights: pd.Series = None): self.returns = returns self.weights = weights def historical_stress_test( self, scenario_name: str, historical_data: pd.DataFrame ) -> Dict[str, float]: """Test portfolio against historical crisis period.""" if scenario_name not in self.HISTORICAL_SCENARIOS: raise ValueError(f"Unknown scenario: {scenario_name}") start, end = self.HISTORICAL_SCENARIOS[scenario_name] # Get returns during crisis crisis_returns = historical_data.loc[start:end] if self.weights is not None: port_returns = (crisis_returns @ self.weights) else: port_returns = crisis_returns total_return = (1 + port_returns).prod() - 1 max_dd = self._calculate_max_dd(port_returns) worst_day = port_returns.min() return { "scenario": scenario_name, "period": f"{start} to {end}", "total_return": total_return, "max_drawdown": max_dd, "worst_day": worst_day, "volatility": port_returns.std() * np.sqrt(252) } def hypothetical_stress_test( self, shocks: Dict[str, float] ) -> float: """ Test portfolio against hypothetical shocks. Args: shocks: Dict of {asset: shock_return} """ if self.weights is None: raise ValueError("Weights required for hypothetical stress test") total_impact = 0 for asset, shock in shocks.items(): if asset in self.weights.index: total_impact += self.weights[asset] * shock return total_impact def monte_carlo_stress( self, n_simulations: int = 10000, horizon_days: int = 21, vol_multiplier: float = 2.0 ) -> Dict[str, float]: """Monte Carlo stress test with elevated volatility.""" mean = self.returns.mean() vol = self.returns.std() * vol_multiplier simulations = np.random.normal( mean, vol, (n_simulations, horizon_days) ) total_returns = (1 + simulations).prod(axis=1) - 1 return { "expected_loss": -total_returns.mean(), "var_95": -np.percentile(total_returns, 5), "var_99": -np.percentile(total_returns, 1), "worst_case": -total_returns.min(), "prob_10pct_loss": (total_returns < -0.10).mean() } def _calculate_max_dd(self, returns: pd.Series) -> float: cumulative = (1 + returns).cumprod() running_max = cumulative.cummax() drawdowns = (cumulative - running_max) / running_max return drawdowns.min() ``` ## Quick Reference ```python # Daily usage metrics = RiskMetrics(returns) print(f"Sharpe: {metrics.sharpe_ratio():.2f}") print(f"Max DD: {metrics.max_drawdown():.2%}") print(f"VaR 95%: {metrics.var_historical(0.95):.2%}") # Full summary summary = metrics.summary() for metric, value in summary.items(): print(f"{metric}: {value:.4f}") ``` ## Best Practices ### Do's - **Use multiple metrics** - No single metric captures all risk - **Consider tail risk** - VaR isn't enough, use CVaR - **Rolling analysis** - Risk changes over time - **Stress test** - Historical and hypothetical - **Document assumptions** - Distribution, lookback, etc. ### Don'ts - **Don't rely on VaR alone** - Underestimates tail risk - **Don't assume normality** - Returns are fat-tailed - **Don't ignore correlation** - Increases in stress - **Don't use short lookbacks** - Miss regime changes - **Don't forget transaction costs** - Affects realized risk ## Resources - [Risk Management and Financial Institutions (John Hull)](https://www.amazon.com/Risk-Management-Financial-Institutions-5th/dp/1119448115) - [Quantitative Risk Management (McNeil, Frey, Embrechts)](https://www.amazon.com/Quantitative-Risk-Management-Techniques-Princeton/dp/0691166277) - [pyfolio Documentation](https://quantopian.github.io/pyfolio/)
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

stride-analysis-patterns

Apply STRIDE methodology to systematically identify threats. Use

security
⭐1
# STRIDE Analysis Patterns Systematic threat identification using the STRIDE methodology. ## When to Use This Skill - Starting new threat modeling sessions - Analyzing existing system architecture - Reviewing security design decisions - Creating threat documentation - Training teams on threat identification - Compliance and audit preparation ## Core Concepts ### 1. STRIDE Categories ``` S - Spoofing → Authentication threats T - Tampering → Integrity threats R - Repudiation → Non-repudiation threats I - Information → Confidentiality threats Disclosure D - Denial of → Availability threats Service E - Elevation of → Authorization threats Privilege ``` ### 2. Threat Analysis Matrix | Category | Question | Control Family | | ------------------- | ----------------------------------------- | -------------- | | **Spoofing** | Can attacker pretend to be someone else? | Authentication | | **Tampering** | Can attacker modify data in transit/rest? | Integrity | | **Repudiation** | Can attacker deny actions? | Logging/Audit | | **Info Disclosure** | Can attacker access unauthorized data? | Encryption | | **DoS** | Can attacker disrupt availability? | Rate limiting | | **Elevation** | Can attacker gain higher privileges? | Authorization | ## Templates ### Template 1: STRIDE Threat Model Document ```markdown # Threat Model: [System Name] ## 1. System Overview ### 1.1 Description [Brief description of the system and its purpose] ### 1.2 Data Flow Diagram ``` [User] --> [Web App] --> [API Gateway] --> [Backend Services] | v [Database] ``` ### 1.3 Trust Boundaries - **External Boundary**: Internet to DMZ - **Internal Boundary**: DMZ to Internal Network - **Data Boundary**: Application to Database ## 2. Assets | Asset | Sensitivity | Description | |-------|-------------|-------------| | User Credentials | High | Authentication tokens, passwords | | Personal Data | High | PII, financial information | | Session Data | Medium | Active user sessions | | Application Logs | Medium | System activity records | | Configuration | High | System settings, secrets | ## 3. STRIDE Analysis ### 3.1 Spoofing Threats | ID | Threat | Target | Impact | Likelihood | |----|--------|--------|--------|------------| | S1 | Session hijacking | User sessions | High | Medium | | S2 | Token forgery | JWT tokens | High | Low | | S3 | Credential stuffing | Login endpoint | High | High | **Mitigations:** - [ ] Implement MFA - [ ] Use secure session management - [ ] Implement account lockout policies ### 3.2 Tampering Threats | ID | Threat | Target | Impact | Likelihood | |----|--------|--------|--------|------------| | T1 | SQL injection | Database queries | Critical | Medium | | T2 | Parameter manipulation | API requests | High | High | | T3 | File upload abuse | File storage | High | Medium | **Mitigations:** - [ ] Input validation on all endpoints - [ ] Parameterized queries - [ ] File type validation ### 3.3 Repudiation Threats | ID | Threat | Target | Impact | Likelihood | |----|--------|--------|--------|------------| | R1 | Transaction denial | Financial ops | High | Medium | | R2 | Access log tampering | Audit logs | Medium | Low | | R3 | Action attribution | User actions | Medium | Medium | **Mitigations:** - [ ] Comprehensive audit logging - [ ] Log integrity protection - [ ] Digital signatures for critical actions ### 3.4 Information Disclosure Threats | ID | Threat | Target | Impact | Likelihood | |----|--------|--------|--------|------------| | I1 | Data breach | User PII | Critical | Medium | | I2 | Error message leakage | System info | Low | High | | I3 | Insecure transmission | Network traffic | High | Medium | **Mitigations:** - [ ] Encryption at rest and in transit - [ ] Sanitize error messages - [ ] Implement TLS 1.3 ### 3.5 Denial of Service Threats | ID | Threat | Target | Impact | Likelihood | |----|--------|--------|--------|------------| | D1 | Resource exhaustion | API servers | High | High | | D2 | Database overload | Database | Critical | Medium | | D3 | Bandwidth saturation | Network | High | Medium | **Mitigations:** - [ ] Rate limiting - [ ] Auto-scaling - [ ] DDoS protection ### 3.6 Elevation of Privilege Threats | ID | Threat | Target | Impact | Likelihood | |----|--------|--------|--------|------------| | E1 | IDOR vulnerabilities | User resources | High | High | | E2 | Role manipulation | Admin access | Critical | Low | | E3 | JWT claim tampering | Authorization | High | Medium | **Mitigations:** - [ ] Proper authorization checks - [ ] Principle of least privilege - [ ] Server-side role validation ## 4. Risk Assessment ### 4.1 Risk Matrix ``` IMPACT Low Med High Crit Low 1 2 3 4 L Med 2 4 6 8 I High 3 6 9 12 K Crit 4 8 12 16 ``` ### 4.2 Prioritized Risks | Rank | Threat | Risk Score | Priority | |------|--------|------------|----------| | 1 | SQL Injection (T1) | 12 | Critical | | 2 | IDOR (E1) | 9 | High | | 3 | Credential Stuffing (S3) | 9 | High | | 4 | Data Breach (I1) | 8 | High | ## 5. Recommendations ### Immediate Actions 1. Implement input validation framework 2. Add rate limiting to authentication endpoints 3. Enable comprehensive audit logging ### Short-term (30 days) 1. Deploy WAF with OWASP ruleset 2. Implement MFA for sensitive operations 3. Encrypt all PII at rest ### Long-term (90 days) 1. Security awareness training 2. Penetration testing 3. Bug bounty program ``` ### Template 2: STRIDE Analysis Code ```python from dataclasses import dataclass, field from enum import Enum from typing import List, Dict, Optional import json class StrideCategory(Enum): SPOOFING = "S" TAMPERING = "T" REPUDIATION = "R" INFORMATION_DISCLOSURE = "I" DENIAL_OF_SERVICE = "D" ELEVATION_OF_PRIVILEGE = "E" class Impact(Enum): LOW = 1 MEDIUM = 2 HIGH = 3 CRITICAL = 4 class Likelihood(Enum): LOW = 1 MEDIUM = 2 HIGH = 3 CRITICAL = 4 @dataclass class Threat: id: str category: StrideCategory title: str description: str target: str impact: Impact likelihood: Likelihood mitigations: List[str] = field(default_factory=list) status: str = "open" @property def risk_score(self) -> int: return self.impact.value * self.likelihood.value @property def risk_level(self) -> str: score = self.risk_score if score >= 12: return "Critical" elif score >= 6: return "High" elif score >= 3: return "Medium" return "Low" @dataclass class Asset: name: str sensitivity: str description: str data_classification: str @dataclass class TrustBoundary: name: str description: str from_zone: str to_zone: str @dataclass class ThreatModel: name: str version: str description: str assets: List[Asset] = field(default_factory=list) boundaries: List[TrustBoundary] = field(default_factory=list) threats: List[Threat] = field(default_factory=list) def add_threat(self, threat: Threat) -> None: self.threats.append(threat) def get_threats_by_category(self, category: StrideCategory) -> List[Threat]: return [t for t in self.threats if t.category == category] def get_critical_threats(self) -> List[Threat]: return [t for t in self.threats if t.risk_level in ("Critical", "High")] def generate_report(self) -> Dict: """Generate threat model report.""" return { "summary": { "name": self.name, "version": self.version, "total_threats": len(self.threats), "critical_threats": len([t for t in self.threats if t.risk_level == "Critical"]), "high_threats": len([t for t in self.threats if t.risk_level == "High"]), }, "by_category": { cat.name: len(self.get_threats_by_category(cat)) for cat in StrideCategory }, "top_risks": [ { "id": t.id, "title": t.title, "risk_score": t.risk_score, "risk_level": t.risk_level } for t in sorted(self.threats, key=lambda x: x.risk_score, reverse=True)[:10] ] } class StrideAnalyzer: """Automated STRIDE analysis helper.""" STRIDE_QUESTIONS = { StrideCategory.SPOOFING: [ "Can an attacker impersonate a legitimate user?", "Are authentication tokens properly validated?", "Can session identifiers be predicted or stolen?", "Is multi-factor authentication available?", ], StrideCategory.TAMPERING: [ "Can data be modified in transit?", "Can data be modified at rest?", "Are input validation controls sufficient?", "Can an attacker manipulate application logic?", ], StrideCategory.REPUDIATION: [ "Are all security-relevant actions logged?", "Can logs be tampered with?", "Is there sufficient attribution for actions?", "Are timestamps reliable and synchronized?", ], StrideCategory.INFORMATION_DISCLOSURE: [ "Is sensitive data encrypted at rest?", "Is sensitive data encrypted in transit?", "Can error messages reveal sensitive information?", "Are access controls properly enforced?", ], StrideCategory.DENIAL_OF_SERVICE: [ "Are rate limits implemented?", "Can resources be exhausted by malicious input?", "Is there protection against amplification attacks?", "Are there single points of failure?", ], StrideCategory.ELEVATION_OF_PRIVILEGE: [ "Are authorization checks performed consistently?", "Can users access other users' resources?", "Can privilege escalation occur through parameter manipulation?", "Is the principle of least privilege followed?", ], } def generate_questionnaire(self, component: str) -> List[Dict]: """Generate STRIDE questionnaire for a component.""" questionnaire = [] for category, questions in self.STRIDE_QUESTIONS.items(): for q in questions: questionnaire.append({ "component": component, "category": category.name, "question": q, "answer": None, "notes": "" }) return questionnaire def suggest_mitigations(self, category: StrideCategory) -> List[str]: """Suggest common mitigations for a STRIDE category.""" mitigations = { StrideCategory.SPOOFING: [ "Implement multi-factor authentication", "Use secure session management", "Implement account lockout policies", "Use cryptographically secure tokens", "Validate authentication at every request", ], StrideCategory.TAMPERING: [ "Implement input validation", "Use parameterized queries", "Apply integrity checks (HMAC, signatures)", "Implement Content Security Policy", "Use immutable infrastructure", ], StrideCategory.REPUDIATION: [ "Enable comprehensive audit logging", "Protect log integrity", "Implement digital signatures", "Use centralized, tamper-evident logging", "Maintain accurate timestamps", ], StrideCategory.INFORMATION_DISCLOSURE: [ "Encrypt data at rest and in transit", "Implement proper access controls", "Sanitize error messages", "Use secure defaults", "Implement data classification", ], StrideCategory.DENIAL_OF_SERVICE: [ "Implement rate limiting", "Use auto-scaling", "Deploy DDoS protection", "Implement circuit breakers", "Set resource quotas", ], StrideCategory.ELEVATION_OF_PRIVILEGE: [ "Implement proper authorization", "Follow principle of least privilege", "Validate permissions server-side", "Use role-based access control", "Implement security boundaries", ], } return mitigations.get(category, []) ``` ### Template 3: Data Flow Diagram Analysis ```python from dataclasses import dataclass from typing import List, Set, Tuple from enum import Enum class ElementType(Enum): EXTERNAL_ENTITY = "external" PROCESS = "process" DATA_STORE = "datastore" DATA_FLOW = "dataflow" @dataclass class DFDElement: id: str name: str type: ElementType trust_level: int # 0 = untrusted, higher = more trusted description: str = "" @dataclass class DataFlow: id: str name: str source: str destination: str data_type: str protocol: str encrypted: bool = False class DFDAnalyzer: """Analyze Data Flow Diagrams for STRIDE threats.""" def __init__(self): self.elements: Dict[str, DFDElement] = {} self.flows: List[DataFlow] = [] def add_element(self, element: DFDElement) -> None: self.elements[element.id] = element def add_flow(self, flow: DataFlow) -> None: self.flows.append(flow) def find_trust_boundary_crossings(self) -> List[Tuple[DataFlow, int]]: """Find data flows that cross trust boundaries.""" crossings = [] for flow in self.flows: source = self.elements.get(flow.source) dest = self.elements.get(flow.destination) if source and dest and source.trust_level != dest.trust_level: trust_diff = abs(source.trust_level - dest.trust_level) crossings.append((flow, trust_diff)) return sorted(crossings, key=lambda x: x[1], reverse=True) def identify_threats_per_element(self) -> Dict[str, List[StrideCategory]]: """Map applicable STRIDE categories to element types.""" threat_mapping = { ElementType.EXTERNAL_ENTITY: [ StrideCategory.SPOOFING, StrideCategory.REPUDIATION, ], ElementType.PROCESS: [ StrideCategory.SPOOFING, StrideCategory.TAMPERING, StrideCategory.REPUDIATION, StrideCategory.INFORMATION_DISCLOSURE, StrideCategory.DENIAL_OF_SERVICE, StrideCategory.ELEVATION_OF_PRIVILEGE, ], ElementType.DATA_STORE: [ StrideCategory.TAMPERING, StrideCategory.REPUDIATION, StrideCategory.INFORMATION_DISCLOSURE, StrideCategory.DENIAL_OF_SERVICE, ], ElementType.DATA_FLOW: [ StrideCategory.TAMPERING, StrideCategory.INFORMATION_DISCLOSURE, StrideCategory.DENIAL_OF_SERVICE, ], } result = {} for elem_id, elem in self.elements.items(): result[elem_id] = threat_mapping.get(elem.type, []) return result def analyze_unencrypted_flows(self) -> List[DataFlow]: """Find unencrypted data flows crossing trust boundaries.""" risky_flows = [] for flow in self.flows: if not flow.encrypted: source = self.elements.get(flow.source) dest = self.elements.get(flow.destination) if source and dest and source.trust_level != dest.trust_level: risky_flows.append(flow) return risky_flows def generate_threat_enumeration(self) -> List[Dict]: """Generate comprehensive threat enumeration.""" threats = [] element_threats = self.identify_threats_per_element() for elem_id, categories in element_threats.items(): elem = self.elements[elem_id] for category in categories: threats.append({ "element_id": elem_id, "element_name": elem.name, "element_type": elem.type.value, "stride_category": category.name, "description": f"{category.name} threat against {elem.name}", "trust_level": elem.trust_level }) return threats ``` ### Template 4: STRIDE per Interaction ```python from typing import List, Dict, Optional from dataclasses import dataclass @dataclass class Interaction: """Represents an interaction between two components.""" id: str source: str target: str action: str data: str protocol: str class StridePerInteraction: """Apply STRIDE to each interaction in the system.""" INTERACTION_THREATS = { # Source type -> Target type -> Applicable threats ("external", "process"): { "S": "External entity spoofing identity to process", "T": "Tampering with data sent to process", "R": "External entity denying sending data", "I": "Data exposure during transmission", "D": "Flooding process with requests", "E": "Exploiting process to gain privileges", }, ("process", "datastore"): { "T": "Process tampering with stored data", "R": "Process denying data modifications", "I": "Unauthorized data access by process", "D": "Process exhausting storage resources", }, ("process", "process"): { "S": "Process spoofing another process", "T": "Tampering with inter-process data", "I": "Data leakage between processes", "D": "One process overwhelming another", "E": "Process gaining elevated access", }, } def analyze_interaction( self, interaction: Interaction, source_type: str, target_type: str ) -> List[Dict]: """Analyze a single interaction for STRIDE threats.""" threats = [] key = (source_type, target_type) applicable_threats = self.INTERACTION_THREATS.get(key, {}) for stride_code, description in applicable_threats.items(): threats.append({ "interaction_id": interaction.id, "source": interaction.source, "target": interaction.target, "stride_category": stride_code, "threat_description": description, "context": f"{interaction.action} - {interaction.data}", }) return threats def generate_threat_matrix( self, interactions: List[Interaction], element_types: Dict[str, str] ) -> List[Dict]:
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

market-sizing-analysis

This skill should be used when the user asks to "calculate TAM",

business
⭐1
# Market Sizing Analysis Comprehensive market sizing methodologies for calculating Total Addressable Market (TAM), Serviceable Available Market (SAM), and Serviceable Obtainable Market (SOM) for startup opportunities. ## Overview Market sizing provides the foundation for startup strategy, fundraising, and business planning. Calculate market opportunity using three complementary methodologies: top-down (industry reports), bottom-up (customer segment calculations), and value theory (willingness to pay). ## Core Concepts ### The Three-Tier Market Framework **TAM (Total Addressable Market)** - Total revenue opportunity if achieving 100% market share - Defines the universe of potential customers - Used for long-term vision and market validation - Example: All email marketing software revenue globally **SAM (Serviceable Available Market)** - Portion of TAM targetable with current product/service - Accounts for geographic, segment, or capability constraints - Represents realistic addressable opportunity - Example: AI-powered email marketing for e-commerce in North America **SOM (Serviceable Obtainable Market)** - Realistic market share achievable in 3-5 years - Accounts for competition, resources, and market dynamics - Used for financial projections and fundraising - Example: 2-5% of SAM based on competitive landscape ### When to Use Each Methodology **Top-Down Analysis** - Use when established market research exists - Best for mature, well-defined markets - Validates market existence and growth - Starts with industry reports and narrows down **Bottom-Up Analysis** - Use when targeting specific customer segments - Best for new or niche markets - Most credible for investors - Builds from customer data and pricing **Value Theory** - Use when creating new market categories - Best for disruptive innovations - Estimates based on value creation - Calculates willingness to pay for problem solution ## Three-Methodology Framework ### Methodology 1: Top-Down Analysis Start with total market size and narrow to addressable segments. **Process:** 1. Identify total market category from research reports 2. Apply geographic filters (target regions) 3. Apply segment filters (target industries/customers) 4. Calculate competitive positioning adjustments **Formula:** ``` TAM = Total Market Category Size SAM = TAM × Geographic % × Segment % SOM = SAM × Realistic Capture Rate (2-5%) ``` **When to use:** Established markets with available research (e.g., SaaS, fintech, e-commerce) **Strengths:** Quick, uses credible data, validates market existence **Limitations:** May overestimate for new categories, less granular ### Methodology 2: Bottom-Up Analysis Build market size from customer segment calculations. **Process:** 1. Define target customer segments 2. Estimate number of potential customers per segment 3. Determine average revenue per customer 4. Calculate realistic penetration rates **Formula:** ``` TAM = Σ (Segment Size × Annual Revenue per Customer) SAM = TAM × (Segments You Can Serve / Total Segments) SOM = SAM × Realistic Penetration Rate (Year 3-5) ``` **When to use:** B2B, niche markets, specific customer segments **Strengths:** Most credible for investors, granular, defensible **Limitations:** Requires detailed customer research, time-intensive ### Methodology 3: Value Theory Calculate based on value created and willingness to pay. **Process:** 1. Identify problem being solved 2. Quantify current cost of problem (time, money, inefficiency) 3. Calculate value of solution (savings, gains, efficiency) 4. Estimate willingness to pay (typically 10-30% of value) 5. Multiply by addressable customer base **Formula:** ``` Value per Customer = Problem Cost × % Solved by Solution Price per Customer = Value × Willingness to Pay % (10-30%) TAM = Total Potential Customers × Price per Customer SAM = TAM × % Meeting Buy Criteria SOM = SAM × Realistic Adoption Rate ``` **When to use:** New categories, disruptive innovations, unclear existing markets **Strengths:** Shows value creation, works for new markets **Limitations:** Requires assumptions, harder to validate ## Step-by-Step Process ### Step 1: Define the Market Clearly specify what market is being measured. **Questions to answer:** - What problem is being solved? - Who are the target customers? - What's the product/service category? - What's the geographic scope? - What's the time horizon? **Example:** - Problem: E-commerce companies struggle with email marketing automation - Customers: E-commerce stores with >$1M annual revenue - Category: AI-powered email marketing software - Geography: North America initially, global expansion - Horizon: 3-5 year opportunity ### Step 2: Gather Data Sources Identify credible data for calculations. **Top-Down Sources:** - Industry research reports (Gartner, Forrester, IDC) - Government statistics (Census, BLS, trade associations) - Public company filings and earnings - Market research firms (Statista, CB Insights, PitchBook) **Bottom-Up Sources:** - Customer interviews and surveys - Sales data and CRM records - Industry databases (LinkedIn, ZoomInfo, Crunchbase) - Competitive intelligence - Academic research **Value Theory Sources:** - Customer problem quantification - Time/cost studies - ROI case studies - Pricing research and willingness-to-pay surveys ### Step 3: Calculate TAM Apply chosen methodology to determine total market. **For Top-Down:** 1. Find total category size from research 2. Document data source and year 3. Apply growth rate if needed 4. Validate with multiple sources **For Bottom-Up:** 1. Count total potential customers 2. Calculate average annual revenue per customer 3. Multiply to get TAM 4. Break down by segment **For Value Theory:** 1. Quantify total addressable customer base 2. Calculate value per customer 3. Estimate pricing based on value 4. Multiply for TAM ### Step 4: Calculate SAM Narrow TAM to serviceable addressable market. **Apply Filters:** - Geographic constraints (regions you can serve) - Product limitations (features you currently have) - Customer requirements (size, industry, use case) - Distribution channel access - Regulatory or compliance restrictions **Formula:** ``` SAM = TAM × (% matching all filters) ``` **Example:** - TAM: $10B global email marketing - Geographic filter: 40% (North America) - Product filter: 30% (e-commerce focus) - Feature filter: 60% (need AI capabilities) - SAM = $10B × 0.40 × 0.30 × 0.60 = $720M ### Step 5: Calculate SOM Determine realistic obtainable market share. **Consider:** - Current market share of competitors - Typical market share for new entrants (2-5%) - Resources available (funding, team, time) - Go-to-market effectiveness - Competitive advantages - Time to achieve (3-5 years typically) **Conservative Approach:** ``` SOM (Year 3) = SAM × 2% SOM (Year 5) = SAM × 5% ``` **Example:** - SAM: $720M - Year 3 SOM: $720M × 2% = $14.4M - Year 5 SOM: $720M × 5% = $36M ### Step 6: Validate and Triangulate Cross-check using multiple methods. **Validation Techniques:** 1. Compare top-down and bottom-up results (should be within 30%) 2. Check against public company revenues in space 3. Validate customer count assumptions 4. Sense-check pricing assumptions 5. Review with industry experts 6. Compare to similar market categories **Red Flags:** - TAM that's too small (< $1B for VC-backed startups) - TAM that's too large (unsupported by data) - SOM that's too aggressive (> 10% in 5 years for new entrant) - Inconsistency between methodologies (> 50% difference) ## Industry-Specific Considerations ### SaaS Markets **Key Metrics:** - Number of potential businesses in target segment - Average contract value (ACV) - Typical market penetration rates - Expansion revenue potential **TAM Calculation:** ``` TAM = Total Target Companies × Average ACV × (1 + Expansion Rate) ``` ### Marketplace Markets **Key Metrics:** - Gross Merchandise Value (GMV) of category - Take rate (% of GMV you capture) - Total transactions or users **TAM Calculation:** ``` TAM = Total Category GMV × Expected Take Rate ``` ### Consumer Markets **Key Metrics:** - Total addressable users/households - Average revenue per user (ARPU) - Engagement frequency **TAM Calculation:** ``` TAM = Total Users × ARPU × Purchase Frequency per Year ``` ### B2B Services **Key Metrics:** - Number of target companies by size/industry - Average project value or retainer - Typical buying frequency **TAM Calculation:** ``` TAM = Total Target Companies × Average Deal Size × Deals per Year ``` ## Presenting Market Sizing ### For Investors **Structure:** 1. Market definition and problem scope 2. TAM/SAM/SOM with methodology 3. Data sources and assumptions 4. Growth projections and drivers 5. Competitive landscape context **Key Points:** - Lead with bottom-up calculation (most credible) - Show triangulation with top-down - Explain conservative assumptions - Link to revenue projections - Highlight market growth rate ### For Strategy **Structure:** 1. Addressable customer segments 2. Prioritization by opportunity size 3. Entry strategy by segment 4. Expected penetration timeline 5. Resource requirements **Key Points:** - Focus on SAM and SOM - Show segment-level detail - Connect to go-to-market plan - Identify expansion opportunities - Discuss competitive positioning ## Common Mistakes to Avoid **Mistake 1: Confusing TAM with SAM** - Don't claim entire market as addressable - Apply realistic product/geographic constraints - Be honest about serviceable market **Mistake 2: Overly Aggressive SOM** - New entrants rarely capture > 5% in 5 years - Account for competition and resources - Show realistic ramp timeline **Mistake 3: Using Only Top-Down** - Investors prefer bottom-up validation - Top-down alone lacks credibility - Always triangulate with multiple methods **Mistake 4: Cherry-Picking Data** - Use consistent, recent data sources - Don't mix methodologies inappropriately - Document all assumptions clearly **Mistake 5: Ignoring Market Dynamics** - Account for market growth/decline - Consider competitive intensity - Factor in switching costs and barriers ## Additional Resources ### Reference Files For detailed methodologies and frameworks: - **`references/methodology-deep-dive.md`** - Comprehensive guide to each methodology with step-by-step worksheets - **`references/data-sources.md`** - Curated list of market research sources, databases, and tools - **`references/industry-templates.md`** - Specific templates for SaaS, marketplace, consumer, B2B, and fintech markets ### Example Files Working examples with complete calculations: - **`examples/saas-market-sizing.md`** - Complete TAM/SAM/SOM for a B2B SaaS product - **`examples/marketplace-sizing.md`** - Marketplace platform market opportunity calculation - **`examples/value-theory-example.md`** - Value-based market sizing for disruptive innovation Use these examples as templates for your own market sizing analysis. Each includes real numbers, data sources, and assumptions documented clearly. ## Quick Start To perform market sizing analysis: 1. **Define the market** - Problem, customers, category, geography 2. **Choose methodology** - Bottom-up (preferred) or top-down + triangulation 3. **Gather data** - Industry reports, customer data, competitive intelligence 4. **Calculate TAM** - Apply methodology formula 5. **Narrow to SAM** - Apply product, geographic, segment filters 6. **Estimate SOM** - 2-5% realistic capture rate 7. **Validate** - Cross-check with alternative methods 8. **Document** - Show methodology, sources, assumptions 9. **Present** - Structure for audience (investors, strategy, operations) For detailed step-by-step guidance on each methodology, reference the files in `references/` directory. For complete worked examples, see `examples/` directory.
👍0
👁️0
🤖 Auto-discovered
🤖system prompt•7 months ago

startup-financial-modeling

This skill should be used when the user asks to "create financial

business
⭐1
# Startup Financial Modeling Build comprehensive 3-5 year financial models with revenue projections, cost structures, cash flow analysis, and scenario planning for early-stage startups. ## Overview Financial modeling provides the quantitative foundation for startup strategy, fundraising, and operational planning. Create realistic projections using cohort-based revenue modeling, detailed cost structures, and scenario analysis to support decision-making and investor presentations. ## Core Components ### Revenue Model **Cohort-Based Projections:** Build revenue from customer acquisition and retention by cohort. **Formula:** ``` MRR = Σ (Cohort Size × Retention Rate × ARPU) ARR = MRR × 12 ``` **Key Inputs:** - Monthly new customer acquisitions - Customer retention rates by month - Average revenue per user (ARPU) - Pricing and packaging assumptions - Expansion revenue (upsells, cross-sells) ### Cost Structure **Operating Expenses Categories:** 1. **Cost of Goods Sold (COGS)** - Hosting and infrastructure - Payment processing fees - Customer support (variable portion) - Third-party services per customer 2. **Sales & Marketing (S&M)** - Customer acquisition cost (CAC) - Marketing programs and advertising - Sales team compensation - Marketing tools and software 3. **Research & Development (R&D)** - Engineering team compensation - Product management - Design and UX - Development tools and infrastructure 4. **General & Administrative (G&A)** - Executive team - Finance, legal, HR - Office and facilities - Insurance and compliance ### Cash Flow Analysis **Components:** - Beginning cash balance - Cash inflows (revenue, fundraising) - Cash outflows (operating expenses, CapEx) - Ending cash balance - Monthly burn rate - Runway (months of cash remaining) **Formula:** ``` Runway = Current Cash Balance / Monthly Burn Rate Monthly Burn = Monthly Revenue - Monthly Expenses ``` ### Headcount Planning **Role-Based Hiring Plan:** Track headcount by department and role. **Key Metrics:** - Fully-loaded cost per employee - Revenue per employee - Headcount by department (% of total) **Typical Ratios (Early-Stage SaaS):** - Engineering: 40-50% - Sales & Marketing: 25-35% - G&A: 10-15% - Customer Success: 5-10% ## Financial Model Structure ### Three-Scenario Framework **Conservative Scenario (P10):** - Slower customer acquisition - Lower pricing or conversion - Higher churn rates - Extended sales cycles - Used for cash management **Base Scenario (P50):** - Most likely outcomes - Realistic assumptions - Primary planning scenario - Used for board reporting **Optimistic Scenario (P90):** - Faster growth - Better unit economics - Lower churn - Used for upside planning ### Time Horizon **Detailed Projections: 3 Years** - Monthly detail for Year 1 - Monthly detail for Year 2 - Quarterly detail for Year 3 **High-Level Projections: Years 4-5** - Annual projections - Key metrics only - Support long-term planning ## Step-by-Step Process ### Step 1: Define Business Model Clarify revenue model and pricing. **SaaS Model:** - Subscription pricing tiers - Annual vs. monthly contracts - Free trial or freemium approach - Expansion revenue strategy **Marketplace Model:** - GMV projections - Take rate (% of transactions) - Buyer and seller economics - Transaction frequency **Transactional Model:** - Transaction volume - Revenue per transaction - Frequency and seasonality ### Step 2: Build Revenue Projections Use cohort-based methodology for accuracy. **Monthly Customer Acquisition:** Define new customers acquired each month. **Retention Curve:** Model customer retention over time. **Typical SaaS Retention:** - Month 1: 100% - Month 3: 90% - Month 6: 85% - Month 12: 75% - Month 24: 70% **Revenue Calculation:** For each cohort, calculate retained customers × ARPU for each month. ### Step 3: Model Cost Structure Break down costs by category and behavior. **Fixed vs. Variable:** - Fixed: Salaries, software, rent - Variable: Hosting, payment processing, support **Scaling Assumptions:** - COGS as % of revenue - S&M as % of revenue (CAC payback) - R&D growth rate - G&A as % of total expenses ### Step 4: Create Hiring Plan Model headcount growth by role and department. **Inputs:** - Starting headcount - Hiring velocity by role - Fully-loaded compensation by role - Benefits and taxes (typically 1.3-1.4x salary) **Example:** ``` Engineer: $150K salary × 1.35 = $202K fully-loaded Sales Rep: $100K OTE × 1.30 = $130K fully-loaded ``` ### Step 5: Project Cash Flow Calculate monthly cash position and runway. **Monthly Cash Flow:** ``` Beginning Cash + Revenue Collected (consider payment terms) - Operating Expenses Paid - CapEx = Ending Cash ``` **Runway Calculation:** ``` If Ending Cash < 0: Funding Need = Negative Cash Balance Runway = 0 Else: Runway = Ending Cash / Average Monthly Burn ``` ### Step 6: Calculate Key Metrics Track metrics that matter for stage. **Revenue Metrics:** - MRR / ARR - Growth rate (MoM, YoY) - Revenue by segment or cohort **Unit Economics:** - CAC (Customer Acquisition Cost) - LTV (Lifetime Value) - CAC Payback Period - LTV / CAC Ratio **Efficiency Metrics:** - Burn multiple (Net Burn / Net New ARR) - Magic number (Net New ARR / S&M Spend) - Rule of 40 (Growth % + Profit Margin %) **Cash Metrics:** - Monthly burn rate - Runway (months) - Cash efficiency ### Step 7: Scenario Analysis Create three scenarios with different assumptions. **Variable Assumptions:** - Customer acquisition rate (±30%) - Churn rate (±20%) - Average contract value (±15%) - CAC (±25%) **Fixed Assumptions:** - Pricing structure - Core operating expenses - Hiring plan (adjust timing, not roles) ## Business Model Templates ### SaaS Financial Model **Revenue Drivers:** - New MRR (customers × ARPU) - Expansion MRR (upsells) - Contraction MRR (downgrades) - Churned MRR (lost customers) **Key Ratios:** - Gross margin: 75-85% - S&M as % revenue: 40-60% (early stage) - CAC payback: < 12 months - Net retention: 100-120% **Example Projection:** ``` Year 1: $500K ARR, 50 customers, $100K MRR by Dec Year 2: $2.5M ARR, 200 customers, $208K MRR by Dec Year 3: $8M ARR, 600 customers, $667K MRR by Dec ``` ### Marketplace Financial Model **Revenue Drivers:** - GMV (Gross Merchandise Value) - Take rate (% of GMV) - Net revenue = GMV × Take rate **Key Ratios:** - Take rate: 10-30% depending on category - CAC for buyers vs. sellers - Contribution margin: 60-70% **Example Projection:** ``` Year 1: $5M GMV, 15% take rate = $750K revenue Year 2: $20M GMV, 15% take rate = $3M revenue Year 3: $60M GMV, 15% take rate = $9M revenue ``` ### E-Commerce Financial Model **Revenue Drivers:** - Traffic (visitors) - Conversion rate - Average order value (AOV) - Purchase frequency **Key Ratios:** - Gross margin: 40-60% - Contribution margin: 20-35% - CAC payback: 3-6 months ### Services / Agency Financial Model **Revenue Drivers:** - Billable hours or projects - Hourly rate or project fee - Utilization rate - Team capacity **Key Ratios:** - Gross margin: 50-70% - Utilization: 70-85% - Revenue per employee ## Fundraising Integration ### Funding Scenario Modeling **Pre-Money Valuation:** Based on metrics and comparables. **Dilution:** ``` Post-Money = Pre-Money + Investment Dilution % = Investment / Post-Money ``` **Use of Funds:** Allocate funding to extend runway and achieve milestones. **Example:** ``` Raise: $5M at $20M pre-money Post-Money: $25M Dilution: 20% Use of Funds: - Product Development: $2M (40%) - Sales & Marketing: $2M (40%) - G&A and Operations: $0.5M (10%) - Working Capital: $0.5M (10%) ``` ### Milestone-Based Planning **Identify Key Milestones:** - Product launch - First $1M ARR - Break-even on CAC - Series A fundraise **Funding Amount:** Ensure runway to achieve next milestone + 6 months buffer. ## Common Pitfalls **Pitfall 1: Overly Optimistic Revenue** - New startups rarely hit aggressive projections - Use conservative customer acquisition assumptions - Model realistic churn rates **Pitfall 2: Underestimating Costs** - Add 20% buffer to expense estimates - Include fully-loaded compensation - Account for software and tools **Pitfall 3: Ignoring Cash Flow Timing** - Revenue ≠ cash (payment terms) - Expenses paid before revenue collected - Model cash conversion carefully **Pitfall 4: Static Headcount** - Hiring takes time (3-6 months to fill roles) - Ramp time for productivity (3-6 months) - Account for attrition (10-15% annually) **Pitfall 5: Not Scenario Planning** - Single scenario is never accurate - Always model conservative case - Plan for what you'll do if base case fails ## Model Validation **Sanity Checks:** - [ ] Revenue growth rate is achievable (3x in Year 2, 2x in Year 3) - [ ] Unit economics are realistic (LTV/CAC > 3, payback < 18 months) - [ ] Burn multiple is reasonable (< 2.0 in Year 2-3) - [ ] Headcount scales with revenue (revenue per employee growing) - [ ] Gross margin is appropriate for business model - [ ] S&M spending aligns with CAC and growth targets **Benchmark Against Peers:** Compare key metrics to similar companies at similar stage. **Investor Feedback:** Share model with advisors or investors for feedback on assumptions. ## Additional Resources ### Reference Files For detailed model structures and advanced techniques: - **`references/model-templates.md`** - Complete financial model templates by business model - **`references/unit-economics.md`** - Deep dive on CAC, LTV, payback, and efficiency metrics - **`references/fundraising-scenarios.md`** - Modeling funding rounds and dilution ### Example Files Working financial models with formulas: - **`examples/saas-financial-model.md`** - Complete 3-year SaaS model with cohort analysis - **`examples/marketplace-model.md`** - Marketplace GMV and take rate projections - **`examples/scenario-analysis.md`** - Three-scenario framework with sensitivities ## Quick Start To create a startup financial model: 1. **Define business model** - Revenue drivers and pricing 2. **Project revenue** - Cohort-based with retention 3. **Model costs** - COGS, S&M, R&D, G&A by month 4. **Plan headcount** - Hiring by role and department 5. **Calculate cash flow** - Revenue - expenses = burn/runway 6. **Compute metrics** - CAC, LTV, burn multiple, runway 7. **Create scenarios** - Conservative, base, optimistic 8. **Validate assumptions** - Sanity check and benchmark 9. **Integrate fundraising** - Model funding rounds and milestones For complete templates and formulas, reference the `references/` and `examples/` files.
👍0
👁️0
🤖 Auto-discovered