Skip to main content
EVOKORE// BROWSE
>

./browse/prompts

11 NODES
šŸ¤–system prompt•7 months ago

team-composition-patterns

Design optimal agent team compositions with sizing heuristics,

coding
⭐1
# Team Composition Patterns Best practices for composing multi-agent teams, selecting team sizes, choosing agent types, and configuring display modes for Claude Code's Agent Teams feature. ## When to Use This Skill - Deciding how many teammates to spawn for a task - Choosing between preset team configurations - Selecting the right agent type (subagent_type) for each role - Configuring teammate display modes (tmux, iTerm2, in-process) - Building custom team compositions for non-standard workflows ## Team Sizing Heuristics | Complexity | Team Size | When to Use | | ------------ | --------- | ----------------------------------------------------------- | | Simple | 1-2 | Single-dimension review, isolated bug, small feature | | Moderate | 2-3 | Multi-file changes, 2-3 concerns, medium features | | Complex | 3-4 | Cross-cutting concerns, large features, deep debugging | | Very Complex | 4-5 | Full-stack features, comprehensive reviews, systemic issues | **Rule of thumb**: Start with the smallest team that covers all required dimensions. Adding teammates increases coordination overhead. ## Preset Team Compositions ### Review Team - **Size**: 3 reviewers - **Agents**: 3x `team-reviewer` - **Default dimensions**: security, performance, architecture - **Use when**: Code changes need multi-dimensional quality assessment ### Debug Team - **Size**: 3 investigators - **Agents**: 3x `team-debugger` - **Default hypotheses**: 3 competing hypotheses - **Use when**: Bug has multiple plausible root causes ### Feature Team - **Size**: 3 (1 lead + 2 implementers) - **Agents**: 1x `team-lead` + 2x `team-implementer` - **Use when**: Feature can be decomposed into parallel work streams ### Fullstack Team - **Size**: 4 (1 lead + 3 implementers) - **Agents**: 1x `team-lead` + 1x frontend `team-implementer` + 1x backend `team-implementer` + 1x test `team-implementer` - **Use when**: Feature spans frontend, backend, and test layers ### Research Team - **Size**: 3 researchers - **Agents**: 3x `general-purpose` - **Default areas**: Each assigned a different research question, module, or topic - **Capabilities**: Codebase search (Grep, Glob, Read), web search (WebSearch, WebFetch) - **Use when**: Need to understand a codebase, research libraries, compare approaches, or gather information from code and web sources in parallel ### Security Team - **Size**: 4 reviewers - **Agents**: 4x `team-reviewer` - **Default dimensions**: OWASP/vulnerabilities, auth/access control, dependencies/supply chain, secrets/configuration - **Use when**: Comprehensive security audit covering multiple attack surfaces ### Migration Team - **Size**: 4 (1 lead + 2 implementers + 1 reviewer) - **Agents**: 1x `team-lead` + 2x `team-implementer` + 1x `team-reviewer` - **Use when**: Large codebase migration (framework upgrade, language port, API version bump) requiring parallel work with correctness verification ## Agent Type Selection When spawning teammates with the Task tool, choose `subagent_type` based on what tools the teammate needs: | Agent Type | Tools Available | Use For | | ------------------------------ | ----------------------------------------- | ---------------------------------------------------------- | | `general-purpose` | All tools (Read, Write, Edit, Bash, etc.) | Implementation, debugging, any task requiring file changes | | `Explore` | Read-only tools (Read, Grep, Glob) | Research, code exploration, analysis | | `Plan` | Read-only tools | Architecture planning, task decomposition | | `agent-teams:team-reviewer` | All tools | Code review with structured findings | | `agent-teams:team-debugger` | All tools | Hypothesis-driven investigation | | `agent-teams:team-implementer` | All tools | Building features within file ownership boundaries | | `agent-teams:team-lead` | All tools | Team orchestration and coordination | **Key distinction**: Read-only agents (Explore, Plan) cannot modify files. Never assign implementation tasks to read-only agents. ## Display Mode Configuration Configure in `~/.claude/settings.json`: ```json { "teammateMode": "tmux" } ``` | Mode | Behavior | Best For | | -------------- | ------------------------------ | ------------------------------------------------- | | `"tmux"` | Each teammate in a tmux pane | Development workflows, monitoring multiple agents | | `"iterm2"` | Each teammate in an iTerm2 tab | macOS users who prefer iTerm2 | | `"in-process"` | All teammates in same process | Simple tasks, CI/CD environments | ## Custom Team Guidelines When building custom teams: 1. **Every team needs a coordinator** — Either designate a `team-lead` or have the user coordinate directly 2. **Match roles to agent types** — Use specialized agents (reviewer, debugger, implementer) when available 3. **Avoid duplicate roles** — Two agents doing the same thing wastes resources 4. **Define boundaries upfront** — Each teammate needs clear ownership of files or responsibilities 5. **Keep it small** — 2-4 teammates is the sweet spot; 5+ requires significant coordination overhead
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

langchain-architecture

Design LLM applications using LangChain 1.x and LangGraph for

coding
⭐1
# LangChain & LangGraph Architecture Master modern LangChain 1.x and LangGraph for building sophisticated LLM applications with agents, state management, memory, and tool integration. ## When to Use This Skill - Building autonomous AI agents with tool access - Implementing complex multi-step LLM workflows - Managing conversation memory and state - Integrating LLMs with external data sources and APIs - Creating modular, reusable LLM application components - Implementing document processing pipelines - Building production-grade LLM applications ## Package Structure (LangChain 1.x) ``` langchain (1.2.x) # High-level orchestration langchain-core (1.2.x) # Core abstractions (messages, prompts, tools) langchain-community # Third-party integrations langgraph # Agent orchestration and state management langchain-openai # OpenAI integrations langchain-anthropic # Anthropic/Claude integrations langchain-voyageai # Voyage AI embeddings langchain-pinecone # Pinecone vector store ``` ## Core Concepts ### 1. LangGraph Agents LangGraph is the standard for building agents in 2026. It provides: **Key Features:** - **StateGraph**: Explicit state management with typed state - **Durable Execution**: Agents persist through failures - **Human-in-the-Loop**: Inspect and modify state at any point - **Memory**: Short-term and long-term memory across sessions - **Checkpointing**: Save and resume agent state **Agent Patterns:** - **ReAct**: Reasoning + Acting with `create_react_agent` - **Plan-and-Execute**: Separate planning and execution nodes - **Multi-Agent**: Supervisor routing between specialized agents - **Tool-Calling**: Structured tool invocation with Pydantic schemas ### 2. State Management LangGraph uses TypedDict for explicit state: ```python from typing import Annotated, TypedDict from langgraph.graph import MessagesState # Simple message-based state class AgentState(MessagesState): """Extends MessagesState with custom fields.""" context: Annotated[list, "retrieved documents"] # Custom state for complex agents class CustomState(TypedDict): messages: Annotated[list, "conversation history"] context: Annotated[dict, "retrieved context"] current_step: str results: list ``` ### 3. Memory Systems Modern memory implementations: - **ConversationBufferMemory**: Stores all messages (short conversations) - **ConversationSummaryMemory**: Summarizes older messages (long conversations) - **ConversationTokenBufferMemory**: Token-based windowing - **VectorStoreRetrieverMemory**: Semantic similarity retrieval - **LangGraph Checkpointers**: Persistent state across sessions ### 4. Document Processing Loading, transforming, and storing documents: **Components:** - **Document Loaders**: Load from various sources - **Text Splitters**: Chunk documents intelligently - **Vector Stores**: Store and retrieve embeddings - **Retrievers**: Fetch relevant documents ### 5. Callbacks & Tracing LangSmith is the standard for observability: - Request/response logging - Token usage tracking - Latency monitoring - Error tracking - Trace visualization ## Quick Start ### Modern ReAct Agent with LangGraph ```python from langgraph.prebuilt import create_react_agent from langgraph.checkpoint.memory import MemorySaver from langchain_anthropic import ChatAnthropic from langchain_core.tools import tool import ast import operator # Initialize LLM (Claude Sonnet 4.6 recommended) llm = ChatAnthropic(model="claude-sonnet-4-6", temperature=0) # Define tools with Pydantic schemas @tool def search_database(query: str) -> str: """Search internal database for information.""" # Your database search logic return f"Results for: {query}" @tool def calculate(expression: str) -> str: """Safely evaluate a mathematical expression. Supports: +, -, *, /, **, %, parentheses Example: '(2 + 3) * 4' returns '20' """ # Safe math evaluation using ast allowed_operators = { ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul, ast.Div: operator.truediv, ast.Pow: operator.pow, ast.Mod: operator.mod, ast.USub: operator.neg, } def _eval(node): if isinstance(node, ast.Constant): return node.value elif isinstance(node, ast.BinOp): left = _eval(node.left) right = _eval(node.right) return allowed_operators[type(node.op)](left, right) elif isinstance(node, ast.UnaryOp): operand = _eval(node.operand) return allowed_operators[type(node.op)](operand) else: raise ValueError(f"Unsupported operation: {type(node)}") try: tree = ast.parse(expression, mode='eval') return str(_eval(tree.body)) except Exception as e: return f"Error: {e}" tools = [search_database, calculate] # Create checkpointer for memory persistence checkpointer = MemorySaver() # Create ReAct agent agent = create_react_agent( llm, tools, checkpointer=checkpointer ) # Run agent with thread ID for memory config = {"configurable": {"thread_id": "user-123"}} result = await agent.ainvoke( {"messages": [("user", "Search for Python tutorials and calculate 25 * 4")]}, config=config ) ``` ## Architecture Patterns ### Pattern 1: RAG with LangGraph ```python from langgraph.graph import StateGraph, START, END from langchain_anthropic import ChatAnthropic from langchain_voyageai import VoyageAIEmbeddings from langchain_pinecone import PineconeVectorStore from langchain_core.documents import Document from langchain_core.prompts import ChatPromptTemplate from typing import TypedDict, Annotated class RAGState(TypedDict): question: str context: Annotated[list[Document], "retrieved documents"] answer: str # Initialize components llm = ChatAnthropic(model="claude-sonnet-4-6") embeddings = VoyageAIEmbeddings(model="voyage-3-large") vectorstore = PineconeVectorStore(index_name="docs", embedding=embeddings) retriever = vectorstore.as_retriever(search_kwargs={"k": 4}) # Define nodes async def retrieve(state: RAGState) -> RAGState: """Retrieve relevant documents.""" docs = await retriever.ainvoke(state["question"]) return {"context": docs} async def generate(state: RAGState) -> RAGState: """Generate answer from context.""" prompt = ChatPromptTemplate.from_template( """Answer based on the context below. If you cannot answer, say so. Context: {context} Question: {question} Answer:""" ) context_text = "\n\n".join(doc.page_content for doc in state["context"]) response = await llm.ainvoke( prompt.format(context=context_text, question=state["question"]) ) return {"answer": response.content} # Build graph builder = StateGraph(RAGState) builder.add_node("retrieve", retrieve) builder.add_node("generate", generate) builder.add_edge(START, "retrieve") builder.add_edge("retrieve", "generate") builder.add_edge("generate", END) rag_chain = builder.compile() # Use the chain result = await rag_chain.ainvoke({"question": "What is the main topic?"}) ``` ### Pattern 2: Custom Agent with Structured Tools ```python from langchain_core.tools import StructuredTool from pydantic import BaseModel, Field class SearchInput(BaseModel): """Input for database search.""" query: str = Field(description="Search query") filters: dict = Field(default={}, description="Optional filters") class EmailInput(BaseModel): """Input for sending email.""" recipient: str = Field(description="Email recipient") subject: str = Field(description="Email subject") content: str = Field(description="Email body") async def search_database(query: str, filters: dict = {}) -> str: """Search internal database for information.""" # Your database search logic return f"Results for '{query}' with filters {filters}" async def send_email(recipient: str, subject: str, content: str) -> str: """Send an email to specified recipient.""" # Email sending logic return f"Email sent to {recipient}" tools = [ StructuredTool.from_function( coroutine=search_database, name="search_database", description="Search internal database", args_schema=SearchInput ), StructuredTool.from_function( coroutine=send_email, name="send_email", description="Send an email", args_schema=EmailInput ) ] agent = create_react_agent(llm, tools) ``` ### Pattern 3: Multi-Step Workflow with StateGraph ```python from langgraph.graph import StateGraph, START, END from typing import TypedDict, Literal class WorkflowState(TypedDict): text: str entities: list analysis: str summary: str current_step: str async def extract_entities(state: WorkflowState) -> WorkflowState: """Extract key entities from text.""" prompt = f"Extract key entities from: {state['text']}\n\nReturn as JSON list." response = await llm.ainvoke(prompt) return {"entities": response.content, "current_step": "analyze"} async def analyze_entities(state: WorkflowState) -> WorkflowState: """Analyze extracted entities.""" prompt = f"Analyze these entities: {state['entities']}\n\nProvide insights." response = await llm.ainvoke(prompt) return {"analysis": response.content, "current_step": "summarize"} async def generate_summary(state: WorkflowState) -> WorkflowState: """Generate final summary.""" prompt = f"""Summarize: Entities: {state['entities']} Analysis: {state['analysis']} Provide a concise summary.""" response = await llm.ainvoke(prompt) return {"summary": response.content, "current_step": "complete"} def route_step(state: WorkflowState) -> Literal["analyze", "summarize", "end"]: """Route to next step based on current state.""" step = state.get("current_step", "extract") if step == "analyze": return "analyze" elif step == "summarize": return "summarize" return "end" # Build workflow builder = StateGraph(WorkflowState) builder.add_node("extract", extract_entities) builder.add_node("analyze", analyze_entities) builder.add_node("summarize", generate_summary) builder.add_edge(START, "extract") builder.add_conditional_edges("extract", route_step, { "analyze": "analyze", "summarize": "summarize", "end": END }) builder.add_conditional_edges("analyze", route_step, { "summarize": "summarize", "end": END }) builder.add_edge("summarize", END) workflow = builder.compile() ``` ### Pattern 4: Multi-Agent Orchestration ```python from langgraph.graph import StateGraph, START, END from langgraph.prebuilt import create_react_agent from langchain_core.messages import HumanMessage from typing import Literal class MultiAgentState(TypedDict): messages: list next_agent: str # Create specialized agents researcher = create_react_agent(llm, research_tools) writer = create_react_agent(llm, writing_tools) reviewer = create_react_agent(llm, review_tools) async def supervisor(state: MultiAgentState) -> MultiAgentState: """Route to appropriate agent based on task.""" prompt = f"""Based on the conversation, which agent should handle this? Options: - researcher: For finding information - writer: For creating content - reviewer: For reviewing and editing - FINISH: Task is complete Messages: {state['messages']} Respond with just the agent name.""" response = await llm.ainvoke(prompt) return {"next_agent": response.content.strip().lower()} def route_to_agent(state: MultiAgentState) -> Literal["researcher", "writer", "reviewer", "end"]: """Route based on supervisor decision.""" next_agent = state.get("next_agent", "").lower() if next_agent == "finish": return "end" return next_agent if next_agent in ["researcher", "writer", "reviewer"] else "end" # Build multi-agent graph builder = StateGraph(MultiAgentState) builder.add_node("supervisor", supervisor) builder.add_node("researcher", researcher) builder.add_node("writer", writer) builder.add_node("reviewer", reviewer) builder.add_edge(START, "supervisor") builder.add_conditional_edges("supervisor", route_to_agent, { "researcher": "researcher", "writer": "writer", "reviewer": "reviewer", "end": END }) # Each agent returns to supervisor for agent in ["researcher", "writer", "reviewer"]: builder.add_edge(agent, "supervisor") multi_agent = builder.compile() ``` ## Memory Management ### Token-Based Memory with LangGraph ```python from langgraph.checkpoint.memory import MemorySaver from langgraph.prebuilt import create_react_agent # In-memory checkpointer (development) checkpointer = MemorySaver() # Create agent with persistent memory agent = create_react_agent(llm, tools, checkpointer=checkpointer) # Each thread_id maintains separate conversation config = {"configurable": {"thread_id": "session-abc123"}} # Messages persist across invocations with same thread_id result1 = await agent.ainvoke({"messages": [("user", "My name is Alice")]}, config) result2 = await agent.ainvoke({"messages": [("user", "What's my name?")]}, config) # Agent remembers: "Your name is Alice" ``` ### Production Memory with PostgreSQL ```python from langgraph.checkpoint.postgres import PostgresSaver # Production checkpointer checkpointer = PostgresSaver.from_conn_string( "postgresql://user:pass@localhost/langgraph" ) agent = create_react_agent(llm, tools, checkpointer=checkpointer) ``` ### Vector Store Memory for Long-Term Context ```python from langchain_community.vectorstores import Chroma from langchain_voyageai import VoyageAIEmbeddings embeddings = VoyageAIEmbeddings(model="voyage-3-large") memory_store = Chroma( collection_name="conversation_memory", embedding_function=embeddings, persist_directory="./memory_db" ) async def retrieve_relevant_memory(query: str, k: int = 5) -> list: """Retrieve relevant past conversations.""" docs = await memory_store.asimilarity_search(query, k=k) return [doc.page_content for doc in docs] async def store_memory(content: str, metadata: dict = {}): """Store conversation in long-term memory.""" await memory_store.aadd_texts([content], metadatas=[metadata]) ``` ## Callback System & LangSmith ### LangSmith Tracing ```python import os from langchain_anthropic import ChatAnthropic # Enable LangSmith tracing os.environ["LANGCHAIN_TRACING_V2"] = "true" os.environ["LANGCHAIN_API_KEY"] = "your-api-key" os.environ["LANGCHAIN_PROJECT"] = "my-project" # All LangChain/LangGraph operations are automatically traced llm = ChatAnthropic(model="claude-sonnet-4-6") ``` ### Custom Callback Handler ```python from langchain_core.callbacks import BaseCallbackHandler from typing import Any, Dict, List class CustomCallbackHandler(BaseCallbackHandler): def on_llm_start( self, serialized: Dict[str, Any], prompts: List[str], **kwargs ) -> None: print(f"LLM started with {len(prompts)} prompts") def on_llm_end(self, response, **kwargs) -> None: print(f"LLM completed: {len(response.generations)} generations") def on_llm_error(self, error: Exception, **kwargs) -> None: print(f"LLM error: {error}") def on_tool_start( self, serialized: Dict[str, Any], input_str: str, **kwargs ) -> None: print(f"Tool started: {serialized.get('name')}") def on_tool_end(self, output: str, **kwargs) -> None: print(f"Tool completed: {output[:100]}...") # Use callbacks result = await agent.ainvoke( {"messages": [("user", "query")]}, config={"callbacks": [CustomCallbackHandler()]} ) ``` ## Streaming Responses ```python from langchain_anthropic import ChatAnthropic llm = ChatAnthropic(model="claude-sonnet-4-6", streaming=True) # Stream tokens async for chunk in llm.astream("Tell me a story"): print(chunk.content, end="", flush=True) # Stream agent events async for event in agent.astream_events( {"messages": [("user", "Search and summarize")]}, version="v2" ): if event["event"] == "on_chat_model_stream": print(event["data"]["chunk"].content, end="") elif event["event"] == "on_tool_start": print(f"\n[Using tool: {event['name']}]") ``` ## Testing Strategies ```python import pytest from unittest.mock import AsyncMock, patch @pytest.mark.asyncio async def test_agent_tool_selection(): """Test agent selects correct tool.""" with patch.object(llm, 'ainvoke') as mock_llm: mock_llm.return_value = AsyncMock(content="Using search_database") result = await agent.ainvoke({ "messages": [("user", "search for documents")] }) # Verify tool was called assert "search_database" in str(result) @pytest.mark.asyncio async def test_memory_persistence(): """Test memory persists across invocations.""" config = {"configurable": {"thread_id": "test-thread"}} # First message await agent.ainvoke( {"messages": [("user", "Remember: the code is 12345")]}, config ) # Second message should remember result = await agent.ainvoke( {"messages": [("user", "What was the code?")]}, config ) assert "12345" in result["messages"][-1].content ``` ## Performance Optimization ### 1. Caching with Redis ```python from langchain_community.cache import RedisCache from langchain_core.globals import set_llm_cache import redis redis_client = redis.Redis.from_url("redis://localhost:6379") set_llm_cache(RedisCache(redis_client)) ``` ### 2. Async Batch Processing ```python import asyncio from langchain_core.documents import Document async def process_documents(documents: list[Document]) -> list: """Process documents in parallel.""" tasks = [process_single(doc) for doc in documents] return await asyncio.gather(*tasks) async def process_single(doc: Document) -> dict: """Process a single document.""" chunks = text_splitter.split_documents([doc]) embeddings = await embeddings_model.aembed_documents( [c.page_content for c in chunks] ) return {"doc_id": doc.metadata.get("id"), "embeddings": embeddings} ``` ### 3. Connection Pooling ```python from langchain_pinecone import PineconeVectorStore from pinecone import Pinecone # Reuse Pinecone client pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"]) index = pc.Index("my-index") # Create vector store with existing index vectorstore = PineconeVectorStore(index=index, embedding=embeddings) ``` ## Resources - [LangChain Documentation](https://python.langchain.com/docs/) - [LangGraph Documentation](https://langchain-ai.github.io/langgraph/) - [LangSmith Platform](https://smith.langchain.com/) - [LangChain GitHub](https://github.com/langchain-ai/langchain) - [LangGraph GitHub](https://github.com/langchain-ai/langgraph) ## Common Pitfalls 1. **Using Deprecated APIs**: Use LangGraph for agents, not `initialize_agent` 2. **Memory Overflow**: Use checkpointers with TTL for long-running agents 3. **Poor Tool Descriptions**: Clear descriptions help LLM select correct tools 4. **Context Window Exceeded**: Use summarization or sliding window memory 5. **No Error Handling**: Wrap too
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

prompt-engineering-patterns

Master advanced prompt engineering techniques to maximize LLM

coding
⭐1
# Prompt Engineering Patterns Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability. ## When to Use This Skill - Designing complex prompts for production LLM applications - Optimizing prompt performance and consistency - Implementing structured reasoning patterns (chain-of-thought, tree-of-thought) - Building few-shot learning systems with dynamic example selection - Creating reusable prompt templates with variable interpolation - Debugging and refining prompts that produce inconsistent outputs - Implementing system prompts for specialized AI assistants - Using structured outputs (JSON mode) for reliable parsing ## Core Capabilities ### 1. Few-Shot Learning - Example selection strategies (semantic similarity, diversity sampling) - Balancing example count with context window constraints - Constructing effective demonstrations with input-output pairs - Dynamic example retrieval from knowledge bases - Handling edge cases through strategic example selection ### 2. Chain-of-Thought Prompting - Step-by-step reasoning elicitation - Zero-shot CoT with "Let's think step by step" - Few-shot CoT with reasoning traces - Self-consistency techniques (sampling multiple reasoning paths) - Verification and validation steps ### 3. Structured Outputs - JSON mode for reliable parsing - Pydantic schema enforcement - Type-safe response handling - Error handling for malformed outputs ### 4. Prompt Optimization - Iterative refinement workflows - A/B testing prompt variations - Measuring prompt performance metrics (accuracy, consistency, latency) - Reducing token usage while maintaining quality - Handling edge cases and failure modes ### 5. Template Systems - Variable interpolation and formatting - Conditional prompt sections - Multi-turn conversation templates - Role-based prompt composition - Modular prompt components ### 6. System Prompt Design - Setting model behavior and constraints - Defining output formats and structure - Establishing role and expertise - Safety guidelines and content policies - Context setting and background information ## Quick Start ```python from langchain_anthropic import ChatAnthropic from langchain_core.prompts import ChatPromptTemplate from pydantic import BaseModel, Field # Define structured output schema class SQLQuery(BaseModel): query: str = Field(description="The SQL query") explanation: str = Field(description="Brief explanation of what the query does") tables_used: list[str] = Field(description="List of tables referenced") # Initialize model with structured output llm = ChatAnthropic(model="claude-sonnet-4-6") structured_llm = llm.with_structured_output(SQLQuery) # Create prompt template prompt = ChatPromptTemplate.from_messages([ ("system", """You are an expert SQL developer. Generate efficient, secure SQL queries. Always use parameterized queries to prevent SQL injection. Explain your reasoning briefly."""), ("user", "Convert this to SQL: {query}") ]) # Create chain chain = prompt | structured_llm # Use result = await chain.ainvoke({ "query": "Find all users who registered in the last 30 days" }) print(result.query) print(result.explanation) ``` ## Key Patterns ### Pattern 1: Structured Output with Pydantic ```python from anthropic import Anthropic from pydantic import BaseModel, Field from typing import Literal import json class SentimentAnalysis(BaseModel): sentiment: Literal["positive", "negative", "neutral"] confidence: float = Field(ge=0, le=1) key_phrases: list[str] reasoning: str async def analyze_sentiment(text: str) -> SentimentAnalysis: """Analyze sentiment with structured output.""" client = Anthropic() message = client.messages.create( model="claude-sonnet-4-6", max_tokens=500, messages=[{ "role": "user", "content": f"""Analyze the sentiment of this text. Text: {text} Respond with JSON matching this schema: {{ "sentiment": "positive" | "negative" | "neutral", "confidence": 0.0-1.0, "key_phrases": ["phrase1", "phrase2"], "reasoning": "brief explanation" }}""" }] ) return SentimentAnalysis(**json.loads(message.content[0].text)) ``` ### Pattern 2: Chain-of-Thought with Self-Verification ```python from langchain_core.prompts import ChatPromptTemplate cot_prompt = ChatPromptTemplate.from_template(""" Solve this problem step by step. Problem: {problem} Instructions: 1. Break down the problem into clear steps 2. Work through each step showing your reasoning 3. State your final answer 4. Verify your answer by checking it against the original problem Format your response as: ## Steps [Your step-by-step reasoning] ## Answer [Your final answer] ## Verification [Check that your answer is correct] """) ``` ### Pattern 3: Few-Shot with Dynamic Example Selection ```python from langchain_voyageai import VoyageAIEmbeddings from langchain_core.example_selectors import SemanticSimilarityExampleSelector from langchain_chroma import Chroma # Create example selector with semantic similarity example_selector = SemanticSimilarityExampleSelector.from_examples( examples=[ {"input": "How do I reset my password?", "output": "Go to Settings > Security > Reset Password"}, {"input": "Where can I see my order history?", "output": "Navigate to Account > Orders"}, {"input": "How do I contact support?", "output": "Click Help > Contact Us or email support@example.com"}, ], embeddings=VoyageAIEmbeddings(model="voyage-3-large"), vectorstore_cls=Chroma, k=2 # Select 2 most similar examples ) async def get_few_shot_prompt(query: str) -> str: """Build prompt with dynamically selected examples.""" examples = await example_selector.aselect_examples({"input": query}) examples_text = "\n".join( f"User: {ex['input']}\nAssistant: {ex['output']}" for ex in examples ) return f"""You are a helpful customer support assistant. Here are some example interactions: {examples_text} Now respond to this query: User: {query} Assistant:""" ``` ### Pattern 4: Progressive Disclosure Start with simple prompts, add complexity only when needed: ```python PROMPT_LEVELS = { # Level 1: Direct instruction "simple": "Summarize this article: {text}", # Level 2: Add constraints "constrained": """Summarize this article in 3 bullet points, focusing on: - Key findings - Main conclusions - Practical implications Article: {text}""", # Level 3: Add reasoning "reasoning": """Read this article carefully. 1. First, identify the main topic and thesis 2. Then, extract the key supporting points 3. Finally, summarize in 3 bullet points Article: {text} Summary:""", # Level 4: Add examples "few_shot": """Read articles and provide concise summaries. Example: Article: "New research shows that regular exercise can reduce anxiety by up to 40%..." Summary: • Regular exercise reduces anxiety by up to 40% • 30 minutes of moderate activity 3x/week is sufficient • Benefits appear within 2 weeks of starting Now summarize this article: Article: {text} Summary:""" } ``` ### Pattern 5: Error Recovery and Fallback ```python from pydantic import BaseModel, ValidationError import json class ResponseWithConfidence(BaseModel): answer: str confidence: float sources: list[str] alternative_interpretations: list[str] = [] ERROR_RECOVERY_PROMPT = """ Answer the question based on the context provided. Context: {context} Question: {question} Instructions: 1. If you can answer confidently (>0.8), provide a direct answer 2. If you're somewhat confident (0.5-0.8), provide your best answer with caveats 3. If you're uncertain (<0.5), explain what information is missing 4. Always provide alternative interpretations if the question is ambiguous Respond in JSON: {{ "answer": "your answer or 'I cannot determine this from the context'", "confidence": 0.0-1.0, "sources": ["relevant context excerpts"], "alternative_interpretations": ["if question is ambiguous"] }} """ async def answer_with_fallback( context: str, question: str, llm ) -> ResponseWithConfidence: """Answer with error recovery and fallback.""" prompt = ERROR_RECOVERY_PROMPT.format(context=context, question=question) try: response = await llm.ainvoke(prompt) return ResponseWithConfidence(**json.loads(response.content)) except (json.JSONDecodeError, ValidationError) as e: # Fallback: try to extract answer without structure simple_prompt = f"Based on: {context}\n\nAnswer: {question}" simple_response = await llm.ainvoke(simple_prompt) return ResponseWithConfidence( answer=simple_response.content, confidence=0.5, sources=["fallback extraction"], alternative_interpretations=[] ) ``` ### Pattern 6: Role-Based System Prompts ```python SYSTEM_PROMPTS = { "analyst": """You are a senior data analyst with expertise in SQL, Python, and business intelligence. Your responsibilities: - Write efficient, well-documented queries - Explain your analysis methodology - Highlight key insights and recommendations - Flag any data quality concerns Communication style: - Be precise and technical when discussing methodology - Translate technical findings into business impact - Use clear visualizations when helpful""", "assistant": """You are a helpful AI assistant focused on accuracy and clarity. Core principles: - Always cite sources when making factual claims - Acknowledge uncertainty rather than guessing - Ask clarifying questions when the request is ambiguous - Provide step-by-step explanations for complex topics Constraints: - Do not provide medical, legal, or financial advice - Redirect harmful requests appropriately - Protect user privacy""", "code_reviewer": """You are a senior software engineer conducting code reviews. Review criteria: - Correctness: Does the code work as intended? - Security: Are there any vulnerabilities? - Performance: Are there efficiency concerns? - Maintainability: Is the code readable and well-structured? - Best practices: Does it follow language idioms? Output format: 1. Summary assessment (approve/request changes) 2. Critical issues (must fix) 3. Suggestions (nice to have) 4. Positive feedback (what's done well)""" } ``` ## Integration Patterns ### With RAG Systems ```python RAG_PROMPT = """You are a knowledgeable assistant that answers questions based on provided context. Context (retrieved from knowledge base): {context} Instructions: 1. Answer ONLY based on the provided context 2. If the context doesn't contain the answer, say "I don't have information about that in my knowledge base" 3. Cite specific passages using [1], [2] notation 4. If the question is ambiguous, ask for clarification Question: {question} Answer:""" ``` ### With Validation and Verification ```python VALIDATED_PROMPT = """Complete the following task: Task: {task} After generating your response, verify it meets ALL these criteria: āœ“ Directly addresses the original request āœ“ Contains no factual errors āœ“ Is appropriately detailed (not too brief, not too verbose) āœ“ Uses proper formatting āœ“ Is safe and appropriate If verification fails on any criterion, revise before responding. Response:""" ``` ## Performance Optimization ### Token Efficiency ```python # Before: Verbose prompt (150+ tokens) verbose_prompt = """ I would like you to please take the following text and provide me with a comprehensive summary of the main points. The summary should capture the key ideas and important details while being concise and easy to understand. """ # After: Concise prompt (30 tokens) concise_prompt = """Summarize the key points concisely: {text} Summary:""" ``` ### Caching Common Prefixes ```python from anthropic import Anthropic client = Anthropic() # Use prompt caching for repeated system prompts response = client.messages.create( model="claude-sonnet-4-6", max_tokens=1000, system=[ { "type": "text", "text": LONG_SYSTEM_PROMPT, "cache_control": {"type": "ephemeral"} } ], messages=[{"role": "user", "content": user_query}] ) ``` ## Best Practices 1. **Be Specific**: Vague prompts produce inconsistent results 2. **Show, Don't Tell**: Examples are more effective than descriptions 3. **Use Structured Outputs**: Enforce schemas with Pydantic for reliability 4. **Test Extensively**: Evaluate on diverse, representative inputs 5. **Iterate Rapidly**: Small changes can have large impacts 6. **Monitor Performance**: Track metrics in production 7. **Version Control**: Treat prompts as code with proper versioning 8. **Document Intent**: Explain why prompts are structured as they are ## Common Pitfalls - **Over-engineering**: Starting with complex prompts before trying simple ones - **Example pollution**: Using examples that don't match the target task - **Context overflow**: Exceeding token limits with excessive examples - **Ambiguous instructions**: Leaving room for multiple interpretations - **Ignoring edge cases**: Not testing on unusual or boundary inputs - **No error handling**: Assuming outputs will always be well-formed - **Hardcoded values**: Not parameterizing prompts for reuse ## Success Metrics Track these KPIs for your prompts: - **Accuracy**: Correctness of outputs - **Consistency**: Reproducibility across similar inputs - **Latency**: Response time (P50, P95, P99) - **Token Usage**: Average tokens per request - **Success Rate**: Percentage of valid, parseable outputs - **User Satisfaction**: Ratings and feedback ## Resources - [Anthropic Prompt Engineering Guide](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering) - [Claude Prompt Caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) - [OpenAI Prompt Engineering](https://platform.openai.com/docs/guides/prompt-engineering) - [LangChain Prompts](https://python.langchain.com/docs/concepts/prompts/)
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

rag-implementation

Build Retrieval-Augmented Generation (RAG) systems for LLM

coding
⭐1
# RAG Implementation Master Retrieval-Augmented Generation (RAG) to build LLM applications that provide accurate, grounded responses using external knowledge sources. ## When to Use This Skill - Building Q&A systems over proprietary documents - Creating chatbots with current, factual information - Implementing semantic search with natural language queries - Reducing hallucinations with grounded responses - Enabling LLMs to access domain-specific knowledge - Building documentation assistants - Creating research tools with source citation ## Core Components ### 1. Vector Databases **Purpose**: Store and retrieve document embeddings efficiently **Options:** - **Pinecone**: Managed, scalable, serverless - **Weaviate**: Open-source, hybrid search, GraphQL - **Milvus**: High performance, on-premise - **Chroma**: Lightweight, easy to use, local development - **Qdrant**: Fast, filtered search, Rust-based - **pgvector**: PostgreSQL extension, SQL integration ### 2. Embeddings **Purpose**: Convert text to numerical vectors for similarity search **Models (2026):** | Model | Dimensions | Best For | |-------|------------|----------| | **voyage-3-large** | 1024 | Claude apps (Anthropic recommended) | | **voyage-code-3** | 1024 | Code search | | **text-embedding-3-large** | 3072 | OpenAI apps, high accuracy | | **text-embedding-3-small** | 1536 | OpenAI apps, cost-effective | | **bge-large-en-v1.5** | 1024 | Open source, local deployment | | **multilingual-e5-large** | 1024 | Multi-language support | ### 3. Retrieval Strategies **Approaches:** - **Dense Retrieval**: Semantic similarity via embeddings - **Sparse Retrieval**: Keyword matching (BM25, TF-IDF) - **Hybrid Search**: Combine dense + sparse with weighted fusion - **Multi-Query**: Generate multiple query variations - **HyDE**: Generate hypothetical documents for better retrieval ### 4. Reranking **Purpose**: Improve retrieval quality by reordering results **Methods:** - **Cross-Encoders**: BERT-based reranking (ms-marco-MiniLM) - **Cohere Rerank**: API-based reranking - **Maximal Marginal Relevance (MMR)**: Diversity + relevance - **LLM-based**: Use LLM to score relevance ## Quick Start with LangGraph ```python from langgraph.graph import StateGraph, START, END from langchain_anthropic import ChatAnthropic from langchain_voyageai import VoyageAIEmbeddings from langchain_pinecone import PineconeVectorStore from langchain_core.documents import Document from langchain_core.prompts import ChatPromptTemplate from langchain_text_splitters import RecursiveCharacterTextSplitter from typing import TypedDict, Annotated class RAGState(TypedDict): question: str context: list[Document] answer: str # Initialize components llm = ChatAnthropic(model="claude-sonnet-4-6") embeddings = VoyageAIEmbeddings(model="voyage-3-large") vectorstore = PineconeVectorStore(index_name="docs", embedding=embeddings) retriever = vectorstore.as_retriever(search_kwargs={"k": 4}) # RAG prompt rag_prompt = ChatPromptTemplate.from_template( """Answer based on the context below. If you cannot answer, say so. Context: {context} Question: {question} Answer:""" ) async def retrieve(state: RAGState) -> RAGState: """Retrieve relevant documents.""" docs = await retriever.ainvoke(state["question"]) return {"context": docs} async def generate(state: RAGState) -> RAGState: """Generate answer from context.""" context_text = "\n\n".join(doc.page_content for doc in state["context"]) messages = rag_prompt.format_messages( context=context_text, question=state["question"] ) response = await llm.ainvoke(messages) return {"answer": response.content} # Build RAG graph builder = StateGraph(RAGState) builder.add_node("retrieve", retrieve) builder.add_node("generate", generate) builder.add_edge(START, "retrieve") builder.add_edge("retrieve", "generate") builder.add_edge("generate", END) rag_chain = builder.compile() # Use result = await rag_chain.ainvoke({"question": "What are the main features?"}) print(result["answer"]) ``` ## Advanced RAG Patterns ### Pattern 1: Hybrid Search with RRF ```python from langchain_community.retrievers import BM25Retriever from langchain.retrievers import EnsembleRetriever # Sparse retriever (BM25 for keyword matching) bm25_retriever = BM25Retriever.from_documents(documents) bm25_retriever.k = 10 # Dense retriever (embeddings for semantic search) dense_retriever = vectorstore.as_retriever(search_kwargs={"k": 10}) # Combine with Reciprocal Rank Fusion weights ensemble_retriever = EnsembleRetriever( retrievers=[bm25_retriever, dense_retriever], weights=[0.3, 0.7] # 30% keyword, 70% semantic ) ``` ### Pattern 2: Multi-Query Retrieval ```python from langchain.retrievers.multi_query import MultiQueryRetriever # Generate multiple query perspectives for better recall multi_query_retriever = MultiQueryRetriever.from_llm( retriever=vectorstore.as_retriever(search_kwargs={"k": 5}), llm=llm ) # Single query → multiple variations → combined results results = await multi_query_retriever.ainvoke("What is the main topic?") ``` ### Pattern 3: Contextual Compression ```python from langchain.retrievers import ContextualCompressionRetriever from langchain.retrievers.document_compressors import LLMChainExtractor # Compressor extracts only relevant portions compressor = LLMChainExtractor.from_llm(llm) compression_retriever = ContextualCompressionRetriever( base_compressor=compressor, base_retriever=vectorstore.as_retriever(search_kwargs={"k": 10}) ) # Returns only relevant parts of documents compressed_docs = await compression_retriever.ainvoke("specific query") ``` ### Pattern 4: Parent Document Retriever ```python from langchain.retrievers import ParentDocumentRetriever from langchain.storage import InMemoryStore from langchain_text_splitters import RecursiveCharacterTextSplitter # Small chunks for precise retrieval, large chunks for context child_splitter = RecursiveCharacterTextSplitter(chunk_size=400, chunk_overlap=50) parent_splitter = RecursiveCharacterTextSplitter(chunk_size=2000, chunk_overlap=200) # Store for parent documents docstore = InMemoryStore() parent_retriever = ParentDocumentRetriever( vectorstore=vectorstore, docstore=docstore, child_splitter=child_splitter, parent_splitter=parent_splitter ) # Add documents (splits children, stores parents) await parent_retriever.aadd_documents(documents) # Retrieval returns parent documents with full context results = await parent_retriever.ainvoke("query") ``` ### Pattern 5: HyDE (Hypothetical Document Embeddings) ```python from langchain_core.prompts import ChatPromptTemplate class HyDEState(TypedDict): question: str hypothetical_doc: str context: list[Document] answer: str hyde_prompt = ChatPromptTemplate.from_template( """Write a detailed passage that would answer this question: Question: {question} Passage:""" ) async def generate_hypothetical(state: HyDEState) -> HyDEState: """Generate hypothetical document for better retrieval.""" messages = hyde_prompt.format_messages(question=state["question"]) response = await llm.ainvoke(messages) return {"hypothetical_doc": response.content} async def retrieve_with_hyde(state: HyDEState) -> HyDEState: """Retrieve using hypothetical document.""" # Use hypothetical doc for retrieval instead of original query docs = await retriever.ainvoke(state["hypothetical_doc"]) return {"context": docs} # Build HyDE RAG graph builder = StateGraph(HyDEState) builder.add_node("hypothetical", generate_hypothetical) builder.add_node("retrieve", retrieve_with_hyde) builder.add_node("generate", generate) builder.add_edge(START, "hypothetical") builder.add_edge("hypothetical", "retrieve") builder.add_edge("retrieve", "generate") builder.add_edge("generate", END) hyde_rag = builder.compile() ``` ## Document Chunking Strategies ### Recursive Character Text Splitter ```python from langchain_text_splitters import RecursiveCharacterTextSplitter splitter = RecursiveCharacterTextSplitter( chunk_size=1000, chunk_overlap=200, length_function=len, separators=["\n\n", "\n", ". ", " ", ""] # Try in order ) chunks = splitter.split_documents(documents) ``` ### Token-Based Splitting ```python from langchain_text_splitters import TokenTextSplitter splitter = TokenTextSplitter( chunk_size=512, chunk_overlap=50, encoding_name="cl100k_base" # OpenAI tiktoken encoding ) ``` ### Semantic Chunking ```python from langchain_experimental.text_splitter import SemanticChunker splitter = SemanticChunker( embeddings=embeddings, breakpoint_threshold_type="percentile", breakpoint_threshold_amount=95 ) ``` ### Markdown Header Splitter ```python from langchain_text_splitters import MarkdownHeaderTextSplitter headers_to_split_on = [ ("#", "Header 1"), ("##", "Header 2"), ("###", "Header 3"), ] splitter = MarkdownHeaderTextSplitter( headers_to_split_on=headers_to_split_on, strip_headers=False ) ``` ## Vector Store Configurations ### Pinecone (Serverless) ```python from pinecone import Pinecone, ServerlessSpec from langchain_pinecone import PineconeVectorStore # Initialize Pinecone client pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"]) # Create index if needed if "my-index" not in pc.list_indexes().names(): pc.create_index( name="my-index", dimension=1024, # voyage-3-large dimensions metric="cosine", spec=ServerlessSpec(cloud="aws", region="us-east-1") ) # Create vector store index = pc.Index("my-index") vectorstore = PineconeVectorStore(index=index, embedding=embeddings) ``` ### Weaviate ```python import weaviate from langchain_weaviate import WeaviateVectorStore client = weaviate.connect_to_local() # or connect_to_weaviate_cloud() vectorstore = WeaviateVectorStore( client=client, index_name="Documents", text_key="content", embedding=embeddings ) ``` ### Chroma (Local Development) ```python from langchain_chroma import Chroma vectorstore = Chroma( collection_name="my_collection", embedding_function=embeddings, persist_directory="./chroma_db" ) ``` ### pgvector (PostgreSQL) ```python from langchain_postgres.vectorstores import PGVector connection_string = "postgresql+psycopg://user:pass@localhost:5432/vectordb" vectorstore = PGVector( embeddings=embeddings, collection_name="documents", connection=connection_string, ) ``` ## Retrieval Optimization ### 1. Metadata Filtering ```python from langchain_core.documents import Document # Add metadata during indexing docs_with_metadata = [] for doc in documents: doc.metadata.update({ "source": doc.metadata.get("source", "unknown"), "category": determine_category(doc.page_content), "date": datetime.now().isoformat() }) docs_with_metadata.append(doc) # Filter during retrieval results = await vectorstore.asimilarity_search( "query", filter={"category": "technical"}, k=5 ) ``` ### 2. Maximal Marginal Relevance (MMR) ```python # Balance relevance with diversity results = await vectorstore.amax_marginal_relevance_search( "query", k=5, fetch_k=20, # Fetch 20, return top 5 diverse lambda_mult=0.5 # 0=max diversity, 1=max relevance ) ``` ### 3. Reranking with Cross-Encoder ```python from sentence_transformers import CrossEncoder reranker = CrossEncoder('cross-encoder/ms-marco-MiniLM-L-6-v2') async def retrieve_and_rerank(query: str, k: int = 5) -> list[Document]: # Get initial results candidates = await vectorstore.asimilarity_search(query, k=20) # Rerank pairs = [[query, doc.page_content] for doc in candidates] scores = reranker.predict(pairs) # Sort by score and take top k ranked = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True) return [doc for doc, score in ranked[:k]] ``` ### 4. Cohere Rerank ```python from langchain.retrievers import CohereRerank from langchain_cohere import CohereRerank reranker = CohereRerank(model="rerank-english-v3.0", top_n=5) # Wrap retriever with reranking reranked_retriever = ContextualCompressionRetriever( base_compressor=reranker, base_retriever=vectorstore.as_retriever(search_kwargs={"k": 20}) ) ``` ## Prompt Engineering for RAG ### Contextual Prompt with Citations ```python rag_prompt = ChatPromptTemplate.from_template( """Answer the question based on the context below. Include citations using [1], [2], etc. If you cannot answer based on the context, say "I don't have enough information." Context: {context} Question: {question} Instructions: 1. Use only information from the context 2. Cite sources with [1], [2] format 3. If uncertain, express uncertainty Answer (with citations):""" ) ``` ### Structured Output for RAG ```python from pydantic import BaseModel, Field class RAGResponse(BaseModel): answer: str = Field(description="The answer based on context") confidence: float = Field(description="Confidence score 0-1") sources: list[str] = Field(description="Source document IDs used") reasoning: str = Field(description="Brief reasoning for the answer") # Use with structured output structured_llm = llm.with_structured_output(RAGResponse) ``` ## Evaluation Metrics ```python from typing import TypedDict class RAGEvalMetrics(TypedDict): retrieval_precision: float # Relevant docs / retrieved docs retrieval_recall: float # Retrieved relevant / total relevant answer_relevance: float # Answer addresses question faithfulness: float # Answer grounded in context context_relevance: float # Context relevant to question async def evaluate_rag_system( rag_chain, test_cases: list[dict] ) -> RAGEvalMetrics: """Evaluate RAG system on test cases.""" metrics = {k: [] for k in RAGEvalMetrics.__annotations__} for test in test_cases: result = await rag_chain.ainvoke({"question": test["question"]}) # Retrieval metrics retrieved_ids = {doc.metadata["id"] for doc in result["context"]} relevant_ids = set(test["relevant_doc_ids"]) precision = len(retrieved_ids & relevant_ids) / len(retrieved_ids) recall = len(retrieved_ids & relevant_ids) / len(relevant_ids) metrics["retrieval_precision"].append(precision) metrics["retrieval_recall"].append(recall) # Use LLM-as-judge for quality metrics quality = await evaluate_answer_quality( question=test["question"], answer=result["answer"], context=result["context"], expected=test.get("expected_answer") ) metrics["answer_relevance"].append(quality["relevance"]) metrics["faithfulness"].append(quality["faithfulness"]) metrics["context_relevance"].append(quality["context_relevance"]) return {k: sum(v) / len(v) for k, v in metrics.items()} ``` ## Resources - [LangChain RAG Tutorial](https://python.langchain.com/docs/tutorials/rag/) - [LangGraph RAG Examples](https://langchain-ai.github.io/langgraph/tutorials/rag/) - [Pinecone Best Practices](https://docs.pinecone.io/guides/get-started/overview) - [Voyage AI Embeddings](https://docs.voyageai.com/) - [RAG Evaluation Guide](https://docs.ragas.io/) ## Best Practices 1. **Chunk Size**: Balance between context (larger) and specificity (smaller) - typically 500-1000 tokens 2. **Overlap**: Use 10-20% overlap to preserve context at boundaries 3. **Metadata**: Include source, page, timestamp for filtering and debugging 4. **Hybrid Search**: Combine semantic and keyword search for best recall 5. **Reranking**: Use cross-encoder reranking for precision-critical applications 6. **Citations**: Always return source documents for transparency 7. **Evaluation**: Continuously test retrieval quality and answer accuracy 8. **Monitoring**: Track retrieval metrics and latency in production ## Common Issues - **Poor Retrieval**: Check embedding quality, chunk size, query formulation - **Irrelevant Results**: Add metadata filtering, use hybrid search, rerank - **Missing Information**: Ensure documents are properly indexed, check chunking - **Slow Queries**: Optimize vector store, use caching, reduce k - **Hallucinations**: Improve grounding prompt, add verification step - **Context Too Long**: Use compression or parent document retriever
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

anti-reversing-techniques

Understand anti-reversing, obfuscation, and protection techniques

security
⭐1
> **AUTHORIZED USE ONLY**: This skill contains dual-use security techniques. Before proceeding with any bypass or analysis: > > 1. **Verify authorization**: Confirm you have explicit written permission from the software owner, or are operating within a legitimate security context (CTF, authorized pentest, malware analysis, security research) > 2. **Document scope**: Ensure your activities fall within the defined scope of your authorization > 3. **Legal compliance**: Understand that unauthorized bypassing of software protection may violate laws (CFAA, DMCA anti-circumvention, etc.) > > **Legitimate use cases**: Malware analysis, authorized penetration testing, CTF competitions, academic security research, analyzing software you own/have rights to # Anti-Reversing Techniques Understanding protection mechanisms encountered during authorized software analysis, security research, and malware analysis. This knowledge helps analysts bypass protections to complete legitimate analysis tasks. ## Anti-Debugging Techniques ### Windows Anti-Debugging #### API-Based Detection ```c // IsDebuggerPresent if (IsDebuggerPresent()) { exit(1); } // CheckRemoteDebuggerPresent BOOL debugged = FALSE; CheckRemoteDebuggerPresent(GetCurrentProcess(), &debugged); if (debugged) exit(1); // NtQueryInformationProcess typedef NTSTATUS (NTAPI *pNtQueryInformationProcess)( HANDLE, PROCESSINFOCLASS, PVOID, ULONG, PULONG); DWORD debugPort = 0; NtQueryInformationProcess( GetCurrentProcess(), ProcessDebugPort, // 7 &debugPort, sizeof(debugPort), NULL ); if (debugPort != 0) exit(1); // Debug flags DWORD debugFlags = 0; NtQueryInformationProcess( GetCurrentProcess(), ProcessDebugFlags, // 0x1F &debugFlags, sizeof(debugFlags), NULL ); if (debugFlags == 0) exit(1); // 0 means being debugged ``` **Bypass Approaches:** ```python # x64dbg: ScyllaHide plugin # Patches common anti-debug checks # Manual patching in debugger: # - Set IsDebuggerPresent return to 0 # - Patch PEB.BeingDebugged to 0 # - Hook NtQueryInformationProcess # IDAPython: Patch checks ida_bytes.patch_byte(check_addr, 0x90) # NOP ``` #### PEB-Based Detection ```c // Direct PEB access #ifdef _WIN64 PPEB peb = (PPEB)__readgsqword(0x60); #else PPEB peb = (PPEB)__readfsdword(0x30); #endif // BeingDebugged flag if (peb->BeingDebugged) exit(1); // NtGlobalFlag // Debugged: 0x70 (FLG_HEAP_ENABLE_TAIL_CHECK | // FLG_HEAP_ENABLE_FREE_CHECK | // FLG_HEAP_VALIDATE_PARAMETERS) if (peb->NtGlobalFlag & 0x70) exit(1); // Heap flags PDWORD heapFlags = (PDWORD)((PBYTE)peb->ProcessHeap + 0x70); if (*heapFlags & 0x50000062) exit(1); ``` **Bypass Approaches:** ```assembly ; In debugger, modify PEB directly ; x64dbg: dump at gs:[60] (x64) or fs:[30] (x86) ; Set BeingDebugged (offset 2) to 0 ; Clear NtGlobalFlag (offset 0xBC for x64) ``` #### Timing-Based Detection ```c // RDTSC timing uint64_t start = __rdtsc(); // ... some code ... uint64_t end = __rdtsc(); if ((end - start) > THRESHOLD) exit(1); // QueryPerformanceCounter LARGE_INTEGER start, end, freq; QueryPerformanceFrequency(&freq); QueryPerformanceCounter(&start); // ... code ... QueryPerformanceCounter(&end); double elapsed = (double)(end.QuadPart - start.QuadPart) / freq.QuadPart; if (elapsed > 0.1) exit(1); // Too slow = debugger // GetTickCount DWORD start = GetTickCount(); // ... code ... if (GetTickCount() - start > 1000) exit(1); ``` **Bypass Approaches:** ``` - Use hardware breakpoints instead of software - Patch timing checks - Use VM with controlled time - Hook timing APIs to return consistent values ``` #### Exception-Based Detection ```c // SEH-based detection __try { __asm { int 3 } // Software breakpoint } __except(EXCEPTION_EXECUTE_HANDLER) { // Normal execution: exception caught return; } // Debugger ate the exception exit(1); // VEH-based detection LONG CALLBACK VectoredHandler(PEXCEPTION_POINTERS ep) { if (ep->ExceptionRecord->ExceptionCode == EXCEPTION_BREAKPOINT) { ep->ContextRecord->Rip++; // Skip INT3 return EXCEPTION_CONTINUE_EXECUTION; } return EXCEPTION_CONTINUE_SEARCH; } ``` ### Linux Anti-Debugging ```c // ptrace self-trace if (ptrace(PTRACE_TRACEME, 0, NULL, NULL) == -1) { // Already being traced exit(1); } // /proc/self/status FILE *f = fopen("/proc/self/status", "r"); char line[256]; while (fgets(line, sizeof(line), f)) { if (strncmp(line, "TracerPid:", 10) == 0) { int tracer_pid = atoi(line + 10); if (tracer_pid != 0) exit(1); } } // Parent process check if (getppid() != 1 && strcmp(get_process_name(getppid()), "bash") != 0) { // Unusual parent (might be debugger) } ``` **Bypass Approaches:** ```bash # LD_PRELOAD to hook ptrace # Compile: gcc -shared -fPIC -o hook.so hook.c long ptrace(int request, ...) { return 0; // Always succeed } # Usage LD_PRELOAD=./hook.so ./target ``` ## Anti-VM Detection ### Hardware Fingerprinting ```c // CPUID-based detection int cpuid_info[4]; __cpuid(cpuid_info, 1); // Check hypervisor bit (bit 31 of ECX) if (cpuid_info[2] & (1 << 31)) { // Running in hypervisor } // CPUID brand string __cpuid(cpuid_info, 0x40000000); char vendor[13] = {0}; memcpy(vendor, &cpuid_info[1], 12); // "VMwareVMware", "Microsoft Hv", "KVMKVMKVM", "VBoxVBoxVBox" // MAC address prefix // VMware: 00:0C:29, 00:50:56 // VirtualBox: 08:00:27 // Hyper-V: 00:15:5D ``` ### Registry/File Detection ```c // Windows registry keys // HKLM\SOFTWARE\VMware, Inc.\VMware Tools // HKLM\SOFTWARE\Oracle\VirtualBox Guest Additions // HKLM\HARDWARE\ACPI\DSDT\VBOX__ // Files // C:\Windows\System32\drivers\vmmouse.sys // C:\Windows\System32\drivers\vmhgfs.sys // C:\Windows\System32\drivers\VBoxMouse.sys // Processes // vmtoolsd.exe, vmwaretray.exe // VBoxService.exe, VBoxTray.exe ``` ### Timing-Based VM Detection ```c // VM exits cause timing anomalies uint64_t start = __rdtsc(); __cpuid(cpuid_info, 0); // Causes VM exit uint64_t end = __rdtsc(); if ((end - start) > 500) { // Likely in VM (CPUID takes longer) } ``` **Bypass Approaches:** ``` - Use bare-metal analysis environment - Harden VM (remove guest tools, change MAC) - Patch detection code - Use specialized analysis VMs (FLARE-VM) ``` ## Code Obfuscation ### Control Flow Obfuscation #### Control Flow Flattening ```c // Original if (cond) { func_a(); } else { func_b(); } func_c(); // Flattened int state = 0; while (1) { switch (state) { case 0: state = cond ? 1 : 2; break; case 1: func_a(); state = 3; break; case 2: func_b(); state = 3; break; case 3: func_c(); return; } } ``` **Analysis Approach:** - Identify state variable - Map state transitions - Reconstruct original flow - Tools: D-810 (IDA), SATURN #### Opaque Predicates ```c // Always true, but complex to analyze int x = rand(); if ((x * x) >= 0) { // Always true real_code(); } else { junk_code(); // Dead code } // Always false if ((x * (x + 1)) % 2 == 1) { // Product of consecutive = even junk_code(); } ``` **Analysis Approach:** - Identify constant expressions - Symbolic execution to prove predicates - Pattern matching for known opaque predicates ### Data Obfuscation #### String Encryption ```c // XOR encryption char decrypt_string(char *enc, int len, char key) { char *dec = malloc(len + 1); for (int i = 0; i < len; i++) { dec[i] = enc[i] ^ key; } dec[len] = 0; return dec; } // Stack strings char url[20]; url[0] = 'h'; url[1] = 't'; url[2] = 't'; url[3] = 'p'; url[4] = ':'; url[5] = '/'; url[6] = '/'; // ... ``` **Analysis Approach:** ```python # FLOSS for automatic string deobfuscation floss malware.exe # IDAPython string decryption def decrypt_xor(ea, length, key): result = "" for i in range(length): byte = ida_bytes.get_byte(ea + i) result += chr(byte ^ key) return result ``` #### API Obfuscation ```c // Dynamic API resolution typedef HANDLE (WINAPI *pCreateFileW)(LPCWSTR, DWORD, DWORD, LPSECURITY_ATTRIBUTES, DWORD, DWORD, HANDLE); HMODULE kernel32 = LoadLibraryA("kernel32.dll"); pCreateFileW myCreateFile = (pCreateFileW)GetProcAddress( kernel32, "CreateFileW"); // API hashing DWORD hash_api(char *name) { DWORD hash = 0; while (*name) { hash = ((hash >> 13) | (hash << 19)) + *name++; } return hash; } // Resolve by hash comparison instead of string ``` **Analysis Approach:** - Identify hash algorithm - Build hash database of known APIs - Use HashDB plugin for IDA - Dynamic analysis to resolve at runtime ### Instruction-Level Obfuscation #### Dead Code Insertion ```asm ; Original mov eax, 1 ; With dead code push ebx ; Dead mov eax, 1 pop ebx ; Dead xor ecx, ecx ; Dead add ecx, ecx ; Dead ``` #### Instruction Substitution ```asm ; Original: xor eax, eax (set to 0) ; Substitutions: sub eax, eax mov eax, 0 and eax, 0 lea eax, [0] ; Original: mov eax, 1 ; Substitutions: xor eax, eax inc eax push 1 pop eax ``` ## Packing and Encryption ### Common Packers ``` UPX - Open source, easy to unpack Themida - Commercial, VM-based protection VMProtect - Commercial, code virtualization ASPack - Compression packer PECompact - Compression packer Enigma - Commercial protector ``` ### Unpacking Methodology ``` 1. Identify packer (DIE, Exeinfo PE, PEiD) 2. Static unpacking (if known packer): - UPX: upx -d packed.exe - Use existing unpackers 3. Dynamic unpacking: a. Find Original Entry Point (OEP) b. Set breakpoint on OEP c. Dump memory when OEP reached d. Fix import table (Scylla, ImpREC) 4. OEP finding techniques: - Hardware breakpoint on stack (ESP trick) - Break on common API calls (GetCommandLineA) - Trace and look for typical entry patterns ``` ### Manual Unpacking Example ``` 1. Load packed binary in x64dbg 2. Note entry point (packer stub) 3. Use ESP trick: - Run to entry - Set hardware breakpoint on [ESP] - Run until breakpoint hits (after PUSHAD/POPAD) 4. Look for JMP to OEP 5. At OEP, use Scylla to: - Dump process - Find imports (IAT autosearch) - Fix dump ``` ## Virtualization-Based Protection ### Code Virtualization ``` Original x86 code is converted to custom bytecode interpreted by embedded VM at runtime. Original: VM Protected: mov eax, 1 push vm_context add eax, 2 call vm_entry ; VM interprets bytecode ; equivalent to original ``` ### Analysis Approaches ``` 1. Identify VM components: - VM entry (dispatcher) - Handler table - Bytecode location - Virtual registers/stack 2. Trace execution: - Log handler calls - Map bytecode to operations - Understand instruction set 3. Lifting/devirtualization: - Map VM instructions back to native - Tools: VMAttack, SATURN, NoVmp 4. Symbolic execution: - Analyze VM semantically - angr, Triton ``` ## Bypass Strategies Summary ### General Principles 1. **Understand the protection**: Identify what technique is used 2. **Find the check**: Locate protection code in binary 3. **Patch or hook**: Modify check to always pass 4. **Use appropriate tools**: ScyllaHide, x64dbg plugins 5. **Document findings**: Keep notes on bypassed protections ### Tool Recommendations ``` Anti-debug bypass: ScyllaHide, TitanHide Unpacking: x64dbg + Scylla, OllyDumpEx Deobfuscation: D-810, SATURN, miasm VM analysis: VMAttack, NoVmp, manual tracing String decryption: FLOSS, custom scripts Symbolic execution: angr, Triton ``` ### Ethical Considerations This knowledge should only be used for: - Authorized security research - Malware analysis (defensive) - CTF competitions - Understanding protections for legitimate purposes - Educational purposes Never use to bypass protections for: - Software piracy - Unauthorized access - Malicious purposes
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

protocol-reverse-engineering

Master network protocol reverse engineering including packet

security
⭐1
# Protocol Reverse Engineering Comprehensive techniques for capturing, analyzing, and documenting network protocols for security research, interoperability, and debugging. ## Traffic Capture ### Wireshark Capture ```bash # Capture on specific interface wireshark -i eth0 -k # Capture with filter wireshark -i eth0 -k -f "port 443" # Capture to file tshark -i eth0 -w capture.pcap # Ring buffer capture (rotate files) tshark -i eth0 -b filesize:100000 -b files:10 -w capture.pcap ``` ### tcpdump Capture ```bash # Basic capture tcpdump -i eth0 -w capture.pcap # With filter tcpdump -i eth0 port 8080 -w capture.pcap # Capture specific bytes tcpdump -i eth0 -s 0 -w capture.pcap # Full packet # Real-time display tcpdump -i eth0 -X port 80 ``` ### Man-in-the-Middle Capture ```bash # mitmproxy for HTTP/HTTPS mitmproxy --mode transparent -p 8080 # SSL/TLS interception mitmproxy --mode transparent --ssl-insecure # Dump to file mitmdump -w traffic.mitm # Burp Suite # Configure browser proxy to 127.0.0.1:8080 ``` ## Protocol Analysis ### Wireshark Analysis ``` # Display filters tcp.port == 8080 http.request.method == "POST" ip.addr == 192.168.1.1 tcp.flags.syn == 1 && tcp.flags.ack == 0 frame contains "password" # Following streams Right-click > Follow > TCP Stream Right-click > Follow > HTTP Stream # Export objects File > Export Objects > HTTP # Decryption Edit > Preferences > Protocols > TLS - (Pre)-Master-Secret log filename - RSA keys list ``` ### tshark Analysis ```bash # Extract specific fields tshark -r capture.pcap -T fields -e ip.src -e ip.dst -e tcp.port # Statistics tshark -r capture.pcap -q -z conv,tcp tshark -r capture.pcap -q -z endpoints,ip # Filter and extract tshark -r capture.pcap -Y "http" -T json > http_traffic.json # Protocol hierarchy tshark -r capture.pcap -q -z io,phs ``` ### Scapy for Custom Analysis ```python from scapy.all import * # Read pcap packets = rdpcap("capture.pcap") # Analyze packets for pkt in packets: if pkt.haslayer(TCP): print(f"Src: {pkt[IP].src}:{pkt[TCP].sport}") print(f"Dst: {pkt[IP].dst}:{pkt[TCP].dport}") if pkt.haslayer(Raw): print(f"Data: {pkt[Raw].load[:50]}") # Filter packets http_packets = [p for p in packets if p.haslayer(TCP) and (p[TCP].sport == 80 or p[TCP].dport == 80)] # Create custom packets pkt = IP(dst="target")/TCP(dport=80)/Raw(load="GET / HTTP/1.1\r\n") send(pkt) ``` ## Protocol Identification ### Common Protocol Signatures ``` HTTP - "HTTP/1." or "GET " or "POST " at start TLS/SSL - 0x16 0x03 (record layer) DNS - UDP port 53, specific header format SMB - 0xFF 0x53 0x4D 0x42 ("SMB" signature) SSH - "SSH-2.0" banner FTP - "220 " response, "USER " command SMTP - "220 " banner, "EHLO" command MySQL - 0x00 length prefix, protocol version PostgreSQL - 0x00 0x00 0x00 startup length Redis - "*" RESP array prefix MongoDB - BSON documents with specific header ``` ### Protocol Header Patterns ``` +--------+--------+--------+--------+ | Magic number / Signature | +--------+--------+--------+--------+ | Version | Flags | +--------+--------+--------+--------+ | Length | Message Type | +--------+--------+--------+--------+ | Sequence Number / Session ID | +--------+--------+--------+--------+ | Payload... | +--------+--------+--------+--------+ ``` ## Binary Protocol Analysis ### Structure Identification ```python # Common patterns in binary protocols # Length-prefixed message struct Message { uint32_t length; # Total message length uint16_t msg_type; # Message type identifier uint8_t flags; # Flags/options uint8_t reserved; # Padding/alignment uint8_t payload[]; # Variable-length payload }; # Type-Length-Value (TLV) struct TLV { uint8_t type; # Field type uint16_t length; # Field length uint8_t value[]; # Field data }; # Fixed header + variable payload struct Packet { uint8_t magic[4]; # "ABCD" signature uint32_t version; uint32_t payload_len; uint32_t checksum; # CRC32 or similar uint8_t payload[]; }; ``` ### Python Protocol Parser ```python import struct from dataclasses import dataclass @dataclass class MessageHeader: magic: bytes version: int msg_type: int length: int @classmethod def from_bytes(cls, data: bytes): magic, version, msg_type, length = struct.unpack( ">4sHHI", data[:12] ) return cls(magic, version, msg_type, length) def parse_messages(data: bytes): offset = 0 messages = [] while offset < len(data): header = MessageHeader.from_bytes(data[offset:]) payload = data[offset+12:offset+12+header.length] messages.append((header, payload)) offset += 12 + header.length return messages # Parse TLV structure def parse_tlv(data: bytes): fields = [] offset = 0 while offset < len(data): field_type = data[offset] length = struct.unpack(">H", data[offset+1:offset+3])[0] value = data[offset+3:offset+3+length] fields.append((field_type, value)) offset += 3 + length return fields ``` ### Hex Dump Analysis ```python def hexdump(data: bytes, width: int = 16): """Format binary data as hex dump.""" lines = [] for i in range(0, len(data), width): chunk = data[i:i+width] hex_part = ' '.join(f'{b:02x}' for b in chunk) ascii_part = ''.join( chr(b) if 32 <= b < 127 else '.' for b in chunk ) lines.append(f'{i:08x} {hex_part:<{width*3}} {ascii_part}') return '\n'.join(lines) # Example output: # 00000000 48 54 54 50 2f 31 2e 31 20 32 30 30 20 4f 4b 0d HTTP/1.1 200 OK. # 00000010 0a 43 6f 6e 74 65 6e 74 2d 54 79 70 65 3a 20 74 .Content-Type: t ``` ## Encryption Analysis ### Identifying Encryption ```python # Entropy analysis - high entropy suggests encryption/compression import math from collections import Counter def entropy(data: bytes) -> float: if not data: return 0.0 counter = Counter(data) probs = [count / len(data) for count in counter.values()] return -sum(p * math.log2(p) for p in probs) # Entropy thresholds: # < 6.0: Likely plaintext or structured data # 6.0-7.5: Possibly compressed # > 7.5: Likely encrypted or random # Common encryption indicators # - High, uniform entropy # - No obvious structure or patterns # - Length often multiple of block size (16 for AES) # - Possible IV at start (16 bytes for AES-CBC) ``` ### TLS Analysis ```bash # Extract TLS metadata tshark -r capture.pcap -Y "ssl.handshake" \ -T fields -e ip.src -e ssl.handshake.ciphersuite # JA3 fingerprinting (client) tshark -r capture.pcap -Y "ssl.handshake.type == 1" \ -T fields -e ssl.handshake.ja3 # JA3S fingerprinting (server) tshark -r capture.pcap -Y "ssl.handshake.type == 2" \ -T fields -e ssl.handshake.ja3s # Certificate extraction tshark -r capture.pcap -Y "ssl.handshake.certificate" \ -T fields -e x509sat.printableString ``` ### Decryption Approaches ```bash # Pre-master secret log (browser) export SSLKEYLOGFILE=/tmp/keys.log # Configure Wireshark # Edit > Preferences > Protocols > TLS # (Pre)-Master-Secret log filename: /tmp/keys.log # Decrypt with private key (if available) # Only works for RSA key exchange # Edit > Preferences > Protocols > TLS > RSA keys list ``` ## Custom Protocol Documentation ### Protocol Specification Template ```markdown # Protocol Name Specification ## Overview Brief description of protocol purpose and design. ## Transport - Layer: TCP/UDP - Port: XXXX - Encryption: TLS 1.2+ ## Message Format ### Header (12 bytes) | Offset | Size | Field | Description | | ------ | ---- | ------- | ----------------------- | | 0 | 4 | Magic | 0x50524F54 ("PROT") | | 4 | 2 | Version | Protocol version (1) | | 6 | 2 | Type | Message type identifier | | 8 | 4 | Length | Payload length in bytes | ### Message Types | Type | Name | Description | | ---- | --------- | ---------------------- | | 0x01 | HELLO | Connection initiation | | 0x02 | HELLO_ACK | Connection accepted | | 0x03 | DATA | Application data | | 0x04 | CLOSE | Connection termination | ### Type 0x01: HELLO | Offset | Size | Field | Description | | ------ | ---- | ---------- | ------------------------ | | 0 | 4 | ClientID | Unique client identifier | | 4 | 2 | Flags | Connection flags | | 6 | var | Extensions | TLV-encoded extensions | ## State Machine ``` [INIT] --HELLO--> [WAIT_ACK] --HELLO_ACK--> [CONNECTED] | DATA/DATA | [CLOSED] <--CLOSE--+ ``` ## Examples ### Connection Establishment ``` Client -> Server: HELLO (ClientID=0x12345678) Server -> Client: HELLO_ACK (Status=OK) Client -> Server: DATA (payload) ``` ``` ### Wireshark Dissector (Lua) ```lua -- custom_protocol.lua local proto = Proto("custom", "Custom Protocol") -- Define fields local f_magic = ProtoField.string("custom.magic", "Magic") local f_version = ProtoField.uint16("custom.version", "Version") local f_type = ProtoField.uint16("custom.type", "Type") local f_length = ProtoField.uint32("custom.length", "Length") local f_payload = ProtoField.bytes("custom.payload", "Payload") proto.fields = { f_magic, f_version, f_type, f_length, f_payload } -- Message type names local msg_types = { [0x01] = "HELLO", [0x02] = "HELLO_ACK", [0x03] = "DATA", [0x04] = "CLOSE" } function proto.dissector(buffer, pinfo, tree) pinfo.cols.protocol = "CUSTOM" local subtree = tree:add(proto, buffer()) -- Parse header subtree:add(f_magic, buffer(0, 4)) subtree:add(f_version, buffer(4, 2)) local msg_type = buffer(6, 2):uint() subtree:add(f_type, buffer(6, 2)):append_text( " (" .. (msg_types[msg_type] or "Unknown") .. ")" ) local length = buffer(8, 4):uint() subtree:add(f_length, buffer(8, 4)) if length > 0 then subtree:add(f_payload, buffer(12, length)) end end -- Register for TCP port local tcp_table = DissectorTable.get("tcp.port") tcp_table:add(8888, proto) ``` ## Active Testing ### Fuzzing with Boofuzz ```python from boofuzz import * def main(): session = Session( target=Target( connection=TCPSocketConnection("target", 8888) ) ) # Define protocol structure s_initialize("HELLO") s_static(b"\x50\x52\x4f\x54") # Magic s_word(1, name="version") # Version s_word(0x01, name="type") # Type (HELLO) s_size("payload", length=4) # Length field s_block_start("payload") s_dword(0x12345678, name="client_id") s_word(0, name="flags") s_block_end() session.connect(s_get("HELLO")) session.fuzz() if __name__ == "__main__": main() ``` ### Replay and Modification ```python from scapy.all import * # Replay captured traffic packets = rdpcap("capture.pcap") for pkt in packets: if pkt.haslayer(TCP) and pkt[TCP].dport == 8888: send(pkt) # Modify and replay for pkt in packets: if pkt.haslayer(Raw): # Modify payload original = pkt[Raw].load modified = original.replace(b"client", b"CLIENT") pkt[Raw].load = modified # Recalculate checksums del pkt[IP].chksum del pkt[TCP].chksum send(pkt) ``` ## Best Practices ### Analysis Workflow 1. **Capture traffic**: Multiple sessions, different scenarios 2. **Identify boundaries**: Message start/end markers 3. **Map structure**: Fixed header, variable payload 4. **Identify fields**: Compare multiple samples 5. **Document format**: Create specification 6. **Validate understanding**: Implement parser/generator 7. **Test edge cases**: Fuzzing, boundary conditions ### Common Patterns to Look For - Magic numbers/signatures at message start - Version fields for compatibility - Length fields (often before variable data) - Type/opcode fields for message identification - Sequence numbers for ordering - Checksums/CRCs for integrity - Timestamps for timing - Session/connection identifiers
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

sast-configuration

Configure Static Application Security Testing (SAST) tools for

security
⭐1
# SAST Configuration Static Application Security Testing (SAST) tool setup, configuration, and custom rule creation for comprehensive security scanning across multiple programming languages. ## Overview This skill provides comprehensive guidance for setting up and configuring SAST tools including Semgrep, SonarQube, and CodeQL. Use this skill when you need to: - Set up SAST scanning in CI/CD pipelines - Create custom security rules for your codebase - Configure quality gates and compliance policies - Optimize scan performance and reduce false positives - Integrate multiple SAST tools for defense-in-depth ## Core Capabilities ### 1. Semgrep Configuration - Custom rule creation with pattern matching - Language-specific security rules (Python, JavaScript, Go, Java, etc.) - CI/CD integration (GitHub Actions, GitLab CI, Jenkins) - False positive tuning and rule optimization - Organizational policy enforcement ### 2. SonarQube Setup - Quality gate configuration - Security hotspot analysis - Code coverage and technical debt tracking - Custom quality profiles for languages - Enterprise integration with LDAP/SAML ### 3. CodeQL Analysis - GitHub Advanced Security integration - Custom query development - Vulnerability variant analysis - Security research workflows - SARIF result processing ## Quick Start ### Initial Assessment 1. Identify primary programming languages in your codebase 2. Determine compliance requirements (PCI-DSS, SOC 2, etc.) 3. Choose SAST tool based on language support and integration needs 4. Review baseline scan to understand current security posture ### Basic Setup ```bash # Semgrep quick start pip install semgrep semgrep --config=auto --error # SonarQube with Docker docker run -d --name sonarqube -p 9000:9000 sonarqube:latest # CodeQL CLI setup gh extension install github/gh-codeql codeql database create mydb --language=python ``` ## Reference Documentation - [Semgrep Rule Creation](references/semgrep-rules.md) - Pattern-based security rule development - [SonarQube Configuration](references/sonarqube-config.md) - Quality gates and profiles - [CodeQL Setup Guide](references/codeql-setup.md) - Query development and workflows ## Templates & Assets - [semgrep-config.yml](assets/semgrep-config.yml) - Production-ready Semgrep configuration - [sonarqube-settings.xml](assets/sonarqube-settings.xml) - SonarQube quality profile template - [run-sast.sh](scripts/run-sast.sh) - Automated SAST execution script ## Integration Patterns ### CI/CD Pipeline Integration ```yaml # GitHub Actions example - name: Run Semgrep uses: returntocorp/semgrep-action@v1 with: config: >- p/security-audit p/owasp-top-ten ``` ### Pre-commit Hook ```bash # .pre-commit-config.yaml - repo: https://github.com/returntocorp/semgrep rev: v1.45.0 hooks: - id: semgrep args: ['--config=auto', '--error'] ``` ## Best Practices 1. **Start with Baseline** - Run initial scan to establish security baseline - Prioritize critical and high severity findings - Create remediation roadmap 2. **Incremental Adoption** - Begin with security-focused rules - Gradually add code quality rules - Implement blocking only for critical issues 3. **False Positive Management** - Document legitimate suppressions - Create allow lists for known safe patterns - Regularly review suppressed findings 4. **Performance Optimization** - Exclude test files and generated code - Use incremental scanning for large codebases - Cache scan results in CI/CD 5. **Team Enablement** - Provide security training for developers - Create internal documentation for common patterns - Establish security champions program ## Common Use Cases ### New Project Setup ```bash ./scripts/run-sast.sh --setup --language python --tools semgrep,sonarqube ``` ### Custom Rule Development ```yaml # See references/semgrep-rules.md for detailed examples rules: - id: hardcoded-jwt-secret pattern: jwt.encode($DATA, "...", ...) message: JWT secret should not be hardcoded severity: ERROR ``` ### Compliance Scanning ```bash # PCI-DSS focused scan semgrep --config p/pci-dss --json -o pci-scan-results.json ``` ## Troubleshooting ### High False Positive Rate - Review and tune rule sensitivity - Add path filters to exclude test files - Use nostmt metadata for noisy patterns - Create organization-specific rule exceptions ### Performance Issues - Enable incremental scanning - Parallelize scans across modules - Optimize rule patterns for efficiency - Cache dependencies and scan results ### Integration Failures - Verify API tokens and credentials - Check network connectivity and proxy settings - Review SARIF output format compatibility - Validate CI/CD runner permissions ## Related Skills - [OWASP Top 10 Checklist](../owasp-top10-checklist/SKILL.md) - [Container Security](../container-security/SKILL.md) - [Dependency Scanning](../dependency-scanning/SKILL.md) ## Tool Comparison | Tool | Best For | Language Support | Cost | Integration | | --------- | ------------------------ | ---------------- | --------------- | ------------- | | Semgrep | Custom rules, fast scans | 30+ languages | Free/Enterprise | Excellent | | SonarQube | Code quality + security | 25+ languages | Free/Commercial | Good | | CodeQL | Deep analysis, research | 10+ languages | Free (OSS) | GitHub native | ## Next Steps 1. Complete initial SAST tool setup 2. Run baseline security scan 3. Create custom rules for organization-specific patterns 4. Integrate into CI/CD pipeline 5. Establish security gate policies 6. Train development team on findings and remediation
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

competitive-landscape

This skill should be used when the user asks to "analyze

business
⭐1
# Competitive Landscape Analysis Comprehensive frameworks for analyzing competition, identifying differentiation opportunities, and developing winning market positioning strategies. ## Overview Understand competitive dynamics using proven frameworks (Porter's Five Forces, Blue Ocean Strategy, positioning maps) to identify opportunities and craft defensible competitive advantages. ## Porter's Five Forces Analyze industry attractiveness and competitive intensity. ### Force 1: Threat of New Entrants **Barriers to Entry:** - Capital requirements - Economies of scale - Switching costs - Brand loyalty - Regulatory barriers - Access to distribution - Network effects **High Threat:** Low barriers, easy to enter (e.g., simple SaaS tools) **Low Threat:** High barriers (e.g., regulated industries, hardware) **Analysis Questions:** - How easy is it for new competitors to enter? - What would it cost to launch a competing product? - Are there network effects or switching costs protecting incumbents? ### Force 2: Bargaining Power of Suppliers **Supplier Power Factors:** - Supplier concentration - Availability of substitutes - Importance to supplier - Switching costs - Forward integration threat **High Power:** Few suppliers, critical inputs (e.g., cloud infrastructure providers) **Low Power:** Many alternatives, commoditized (e.g., generic services) **Analysis Questions:** - Who are our critical suppliers? - Could they raise prices or reduce quality? - Can we switch suppliers easily? ### Force 3: Bargaining Power of Buyers **Buyer Power Factors:** - Buyer concentration - Volume purchased - Product differentiation - Price sensitivity - Backward integration threat **High Power:** Few large customers, standardized products (e.g., enterprise deals) **Low Power:** Many small customers, differentiated product (e.g., consumer subscriptions) **Analysis Questions:** - Can customers easily switch to competitors? - Do few customers generate most revenue? - How price-sensitive are buyers? ### Force 4: Threat of Substitutes **Substitute Considerations:** - Alternative solutions - Price-performance tradeoff - Switching costs - Buyer propensity to substitute **High Threat:** Many alternatives, low switching cost (e.g., productivity software) **Low Threat:** Unique solution, high switching cost (e.g., ERP systems) **Analysis Questions:** - What alternative ways can customers solve this problem? - How do substitutes compare on price and performance? - What's the cost to switch to a substitute? ### Force 5: Competitive Rivalry **Rivalry Intensity Factors:** - Number of competitors - Industry growth rate - Product differentiation - Exit barriers - Strategic stakes **High Rivalry:** Many competitors, slow growth, commoditized (e.g., email marketing) **Low Rivalry:** Few competitors, fast growth, differentiated (e.g., emerging AI tools) **Analysis Questions:** - How many direct competitors exist? - Is the market growing or stagnant? - How differentiated are offerings? - Are competitors competing on price or value? ### Forces Analysis Summary Create a scorecard: | Force | Intensity (1-5) | Impact | Key Factors | | -------------- | --------------- | ------ | --------------------------------- | | New Entrants | 3 | Medium | Low barriers but network effects | | Supplier Power | 2 | Low | Many cloud providers | | Buyer Power | 4 | High | Enterprise customers concentrated | | Substitutes | 3 | Medium | Manual processes alternative | | Rivalry | 4 | High | 10+ direct competitors | **Overall Assessment:** Moderate industry attractiveness with high rivalry and buyer power ## Blue Ocean Strategy Identify uncontested market space through value innovation. ### Four Actions Framework **Eliminate:** What factors can be eliminated that the industry takes for granted? **Reduce:** What factors can be reduced well below industry standard? **Raise:** What factors can be raised well above industry standard? **Create:** What factors can be created that the industry never offered? ### Strategy Canvas Map your offering vs. competitors on key factors. **Example: Budget Hotels** ``` High | ā˜… Traditional Hotels | ā˜… Budget Hotels (new) | Low |___________________________________ Price Luxury Convenience Cleanliness Budget Hotel Strategy: - Eliminate: Luxury amenities, room service - Reduce: Lobby size, staff - Raise: Cleanliness, online booking - Create: Self-service kiosks, mobile app ``` ### Value Innovation Find the sweet spot: Lower cost + higher value **Steps:** 1. Map industry competing factors 2. Identify factors to eliminate/reduce (cost savings) 3. Identify factors to raise/create (differentiation) 4. Validate that combination creates new market space ## Competitive Positioning ### Positioning Map Plot competitors on 2-3 key dimensions. **Example Dimensions:** - Price vs. Features - Complexity vs. Ease of Use - Enterprise vs. SMB Focus - Self-Service vs. High-Touch - Generalist vs. Specialist **How to Create:** 1. Choose 2 dimensions most important to customers 2. Plot all competitors 3. Identify gaps (white space) 4. Validate gap represents real customer need **Example:** ``` High Price | | ā˜… Enterprise A ā˜… Enterprise B | | ā— Our Position (gap) | | ā˜… Competitor C ā˜… Competitor D | Low Price |____________________________________________ Simple Complex ``` ### Differentiation Strategy **How to Differentiate:** 1. **Product Differentiation** - Unique features - Superior performance - Better design/UX - Integration ecosystem 2. **Service Differentiation** - Customer support quality - Onboarding experience - Response time - Success programs 3. **Brand Differentiation** - Trust and reputation - Thought leadership - Community - Values alignment 4. **Price Differentiation** - Premium positioning - Value positioning - Transparent pricing - Flexible packaging ### Positioning Statement Framework ``` For [target customer] Who [statement of need or opportunity] Our product is [product category] That [statement of key benefit] Unlike [primary competitive alternative] Our product [statement of primary differentiation] ``` **Example:** ``` For e-commerce companies Who struggle with email marketing automation Our product is an AI-powered email platform That increases conversion rates by 40% Unlike Klaviyo and Mailchimp Our product uses AI to personalize at scale ``` ## Competitive Intelligence ### Information Gathering **Public Sources:** - Company websites and blogs - Press releases and news - Job postings (hint at strategy) - Customer reviews (G2, Capterra) - Social media and forums - Glassdoor (employee insights) - SEC filings (public companies) - Patent filings **Direct Research:** - Customer interviews - Win/loss analysis - Sales team feedback - Product demos and trials - Conference attendance ### Competitor Profile Template For each key competitor, document: **Company Overview:** - Founded, HQ, funding, size - Leadership team - Company stage and trajectory **Product:** - Core features - Target customers - Pricing and packaging - Technology stack - Recent launches **Go-to-Market:** - Sales model (self-serve, sales-led) - Marketing strategy - Distribution channels - Partnerships **Strengths:** - What they do better than anyone - Key competitive advantages - Market position **Weaknesses:** - Gaps in product - Customer complaints - Operational challenges **Strategy:** - Stated direction - Inferred priorities - Likely next moves ## Competitive Pricing Analysis ### Price Positioning **Premium (Top 25%):** - Superior product/service - Strong brand - High-touch sales - Enterprise focus **Mid-Market (Middle 50%):** - Balanced value - Standard features - Mixed sales model - Broad market **Value (Bottom 25%):** - Basic functionality - Self-service - Cost leadership - High volume, low margin ### Pricing Comparison Matrix | Competitor | Entry Price | Mid Tier | Enterprise | Model | | ------------ | ----------- | -------- | ---------- | ------------ | | Competitor A | $29/mo | $99/mo | Custom | Subscription | | Competitor B | $49/mo | $199/mo | $499/mo | Subscription | | Us | $39/mo | $129/mo | Custom | Subscription | **Analysis:** - Are we priced competitively? - What does our pricing signal? - Are there gaps in our packaging? ## Go-to-Market Strategy ### Market Entry Strategies **Direct Competition:** - Head-to-head against established players - Requires differentiation and resources - Example: Better features at lower price **Niche Focus:** - Target underserved segment - Become specialist vs. generalist - Example: "Salesforce for real estate" **Disruptive Innovation:** - Target non-consumers or low end - Improve over time to move upmarket - Example: Freemium model disrupting enterprise **Platform Play:** - Build ecosystem and network effects - Aggregate complementary services - Example: Marketplace or API platform ### Beachhead Market **Characteristics of Good Beachhead:** - Specific, reachable segment - Acute pain you solve well - Limited competition - Willing to pay - Can lead to expansion **Example:** Instead of "project management software", target "project management for construction teams" ## Competitive Advantage ### Sustainable Advantages **Network Effects:** - Value increases with users - Example: Slack, marketplaces **Switching Costs:** - High cost to change - Example: CRM systems with data **Economies of Scale:** - Unit costs decrease with volume - Example: Cloud infrastructure **Brand:** - Trust and reputation - Example: Security software **Proprietary Technology:** - Patents or trade secrets - Example: Algorithms, data **Regulatory:** - Licenses or approvals - Example: Fintech, healthcare ### Testing Your Advantage Ask: - Can competitors copy this in < 2 years? - Does this matter to customers? - Do we execute this better than anyone? - Is this advantage durable? If "no" to any, it's not a sustainable advantage. ## Competitive Monitoring ### What to Track **Product Changes:** - New features - Pricing changes - Packaging adjustments **Market Signals:** - Funding announcements - Key hires (especially leadership) - Customer wins/losses - Partnerships **Performance Metrics:** - Revenue (if public or disclosed) - Customer count - Growth rate - Market share estimates ### Monitoring Cadence **Weekly:** - Product release notes - News mentions **Monthly:** - Win/loss analysis review - Positioning map updates **Quarterly:** - Deep competitive review - Strategy adjustment **Annually:** - Major strategy reassessment - Market trends analysis ## Additional Resources ### Reference Files - **`references/frameworks-deep-dive.md`** - Detailed application of each framework with worksheets - **`references/intel-sources.md`** - Comprehensive list of competitive intelligence sources ### Example Files - **`examples/competitor-analysis.md`** - Complete competitive analysis for a SaaS startup - **`examples/positioning-workshop.md`** - Step-by-step positioning development process ## Quick Start To analyze competitive landscape: 1. **Identify competitors** - Direct, indirect, and future threats 2. **Apply Porter's Five Forces** - Assess industry attractiveness 3. **Create positioning map** - Visualize competitive space 4. **Profile top 3-5 competitors** - Deep dive on key rivals 5. **Identify differentiation** - What makes you unique 6. **Analyze pricing** - Where do you fit? 7. **Assess advantages** - What's defensible? 8. **Develop strategy** - How to win For detailed frameworks and examples, see `references/` and `examples/`.
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

market-sizing-analysis

This skill should be used when the user asks to "calculate TAM",

business
⭐1
# Market Sizing Analysis Comprehensive market sizing methodologies for calculating Total Addressable Market (TAM), Serviceable Available Market (SAM), and Serviceable Obtainable Market (SOM) for startup opportunities. ## Overview Market sizing provides the foundation for startup strategy, fundraising, and business planning. Calculate market opportunity using three complementary methodologies: top-down (industry reports), bottom-up (customer segment calculations), and value theory (willingness to pay). ## Core Concepts ### The Three-Tier Market Framework **TAM (Total Addressable Market)** - Total revenue opportunity if achieving 100% market share - Defines the universe of potential customers - Used for long-term vision and market validation - Example: All email marketing software revenue globally **SAM (Serviceable Available Market)** - Portion of TAM targetable with current product/service - Accounts for geographic, segment, or capability constraints - Represents realistic addressable opportunity - Example: AI-powered email marketing for e-commerce in North America **SOM (Serviceable Obtainable Market)** - Realistic market share achievable in 3-5 years - Accounts for competition, resources, and market dynamics - Used for financial projections and fundraising - Example: 2-5% of SAM based on competitive landscape ### When to Use Each Methodology **Top-Down Analysis** - Use when established market research exists - Best for mature, well-defined markets - Validates market existence and growth - Starts with industry reports and narrows down **Bottom-Up Analysis** - Use when targeting specific customer segments - Best for new or niche markets - Most credible for investors - Builds from customer data and pricing **Value Theory** - Use when creating new market categories - Best for disruptive innovations - Estimates based on value creation - Calculates willingness to pay for problem solution ## Three-Methodology Framework ### Methodology 1: Top-Down Analysis Start with total market size and narrow to addressable segments. **Process:** 1. Identify total market category from research reports 2. Apply geographic filters (target regions) 3. Apply segment filters (target industries/customers) 4. Calculate competitive positioning adjustments **Formula:** ``` TAM = Total Market Category Size SAM = TAM Ɨ Geographic % Ɨ Segment % SOM = SAM Ɨ Realistic Capture Rate (2-5%) ``` **When to use:** Established markets with available research (e.g., SaaS, fintech, e-commerce) **Strengths:** Quick, uses credible data, validates market existence **Limitations:** May overestimate for new categories, less granular ### Methodology 2: Bottom-Up Analysis Build market size from customer segment calculations. **Process:** 1. Define target customer segments 2. Estimate number of potential customers per segment 3. Determine average revenue per customer 4. Calculate realistic penetration rates **Formula:** ``` TAM = Ī£ (Segment Size Ɨ Annual Revenue per Customer) SAM = TAM Ɨ (Segments You Can Serve / Total Segments) SOM = SAM Ɨ Realistic Penetration Rate (Year 3-5) ``` **When to use:** B2B, niche markets, specific customer segments **Strengths:** Most credible for investors, granular, defensible **Limitations:** Requires detailed customer research, time-intensive ### Methodology 3: Value Theory Calculate based on value created and willingness to pay. **Process:** 1. Identify problem being solved 2. Quantify current cost of problem (time, money, inefficiency) 3. Calculate value of solution (savings, gains, efficiency) 4. Estimate willingness to pay (typically 10-30% of value) 5. Multiply by addressable customer base **Formula:** ``` Value per Customer = Problem Cost Ɨ % Solved by Solution Price per Customer = Value Ɨ Willingness to Pay % (10-30%) TAM = Total Potential Customers Ɨ Price per Customer SAM = TAM Ɨ % Meeting Buy Criteria SOM = SAM Ɨ Realistic Adoption Rate ``` **When to use:** New categories, disruptive innovations, unclear existing markets **Strengths:** Shows value creation, works for new markets **Limitations:** Requires assumptions, harder to validate ## Step-by-Step Process ### Step 1: Define the Market Clearly specify what market is being measured. **Questions to answer:** - What problem is being solved? - Who are the target customers? - What's the product/service category? - What's the geographic scope? - What's the time horizon? **Example:** - Problem: E-commerce companies struggle with email marketing automation - Customers: E-commerce stores with >$1M annual revenue - Category: AI-powered email marketing software - Geography: North America initially, global expansion - Horizon: 3-5 year opportunity ### Step 2: Gather Data Sources Identify credible data for calculations. **Top-Down Sources:** - Industry research reports (Gartner, Forrester, IDC) - Government statistics (Census, BLS, trade associations) - Public company filings and earnings - Market research firms (Statista, CB Insights, PitchBook) **Bottom-Up Sources:** - Customer interviews and surveys - Sales data and CRM records - Industry databases (LinkedIn, ZoomInfo, Crunchbase) - Competitive intelligence - Academic research **Value Theory Sources:** - Customer problem quantification - Time/cost studies - ROI case studies - Pricing research and willingness-to-pay surveys ### Step 3: Calculate TAM Apply chosen methodology to determine total market. **For Top-Down:** 1. Find total category size from research 2. Document data source and year 3. Apply growth rate if needed 4. Validate with multiple sources **For Bottom-Up:** 1. Count total potential customers 2. Calculate average annual revenue per customer 3. Multiply to get TAM 4. Break down by segment **For Value Theory:** 1. Quantify total addressable customer base 2. Calculate value per customer 3. Estimate pricing based on value 4. Multiply for TAM ### Step 4: Calculate SAM Narrow TAM to serviceable addressable market. **Apply Filters:** - Geographic constraints (regions you can serve) - Product limitations (features you currently have) - Customer requirements (size, industry, use case) - Distribution channel access - Regulatory or compliance restrictions **Formula:** ``` SAM = TAM Ɨ (% matching all filters) ``` **Example:** - TAM: $10B global email marketing - Geographic filter: 40% (North America) - Product filter: 30% (e-commerce focus) - Feature filter: 60% (need AI capabilities) - SAM = $10B Ɨ 0.40 Ɨ 0.30 Ɨ 0.60 = $720M ### Step 5: Calculate SOM Determine realistic obtainable market share. **Consider:** - Current market share of competitors - Typical market share for new entrants (2-5%) - Resources available (funding, team, time) - Go-to-market effectiveness - Competitive advantages - Time to achieve (3-5 years typically) **Conservative Approach:** ``` SOM (Year 3) = SAM Ɨ 2% SOM (Year 5) = SAM Ɨ 5% ``` **Example:** - SAM: $720M - Year 3 SOM: $720M Ɨ 2% = $14.4M - Year 5 SOM: $720M Ɨ 5% = $36M ### Step 6: Validate and Triangulate Cross-check using multiple methods. **Validation Techniques:** 1. Compare top-down and bottom-up results (should be within 30%) 2. Check against public company revenues in space 3. Validate customer count assumptions 4. Sense-check pricing assumptions 5. Review with industry experts 6. Compare to similar market categories **Red Flags:** - TAM that's too small (< $1B for VC-backed startups) - TAM that's too large (unsupported by data) - SOM that's too aggressive (> 10% in 5 years for new entrant) - Inconsistency between methodologies (> 50% difference) ## Industry-Specific Considerations ### SaaS Markets **Key Metrics:** - Number of potential businesses in target segment - Average contract value (ACV) - Typical market penetration rates - Expansion revenue potential **TAM Calculation:** ``` TAM = Total Target Companies Ɨ Average ACV Ɨ (1 + Expansion Rate) ``` ### Marketplace Markets **Key Metrics:** - Gross Merchandise Value (GMV) of category - Take rate (% of GMV you capture) - Total transactions or users **TAM Calculation:** ``` TAM = Total Category GMV Ɨ Expected Take Rate ``` ### Consumer Markets **Key Metrics:** - Total addressable users/households - Average revenue per user (ARPU) - Engagement frequency **TAM Calculation:** ``` TAM = Total Users Ɨ ARPU Ɨ Purchase Frequency per Year ``` ### B2B Services **Key Metrics:** - Number of target companies by size/industry - Average project value or retainer - Typical buying frequency **TAM Calculation:** ``` TAM = Total Target Companies Ɨ Average Deal Size Ɨ Deals per Year ``` ## Presenting Market Sizing ### For Investors **Structure:** 1. Market definition and problem scope 2. TAM/SAM/SOM with methodology 3. Data sources and assumptions 4. Growth projections and drivers 5. Competitive landscape context **Key Points:** - Lead with bottom-up calculation (most credible) - Show triangulation with top-down - Explain conservative assumptions - Link to revenue projections - Highlight market growth rate ### For Strategy **Structure:** 1. Addressable customer segments 2. Prioritization by opportunity size 3. Entry strategy by segment 4. Expected penetration timeline 5. Resource requirements **Key Points:** - Focus on SAM and SOM - Show segment-level detail - Connect to go-to-market plan - Identify expansion opportunities - Discuss competitive positioning ## Common Mistakes to Avoid **Mistake 1: Confusing TAM with SAM** - Don't claim entire market as addressable - Apply realistic product/geographic constraints - Be honest about serviceable market **Mistake 2: Overly Aggressive SOM** - New entrants rarely capture > 5% in 5 years - Account for competition and resources - Show realistic ramp timeline **Mistake 3: Using Only Top-Down** - Investors prefer bottom-up validation - Top-down alone lacks credibility - Always triangulate with multiple methods **Mistake 4: Cherry-Picking Data** - Use consistent, recent data sources - Don't mix methodologies inappropriately - Document all assumptions clearly **Mistake 5: Ignoring Market Dynamics** - Account for market growth/decline - Consider competitive intensity - Factor in switching costs and barriers ## Additional Resources ### Reference Files For detailed methodologies and frameworks: - **`references/methodology-deep-dive.md`** - Comprehensive guide to each methodology with step-by-step worksheets - **`references/data-sources.md`** - Curated list of market research sources, databases, and tools - **`references/industry-templates.md`** - Specific templates for SaaS, marketplace, consumer, B2B, and fintech markets ### Example Files Working examples with complete calculations: - **`examples/saas-market-sizing.md`** - Complete TAM/SAM/SOM for a B2B SaaS product - **`examples/marketplace-sizing.md`** - Marketplace platform market opportunity calculation - **`examples/value-theory-example.md`** - Value-based market sizing for disruptive innovation Use these examples as templates for your own market sizing analysis. Each includes real numbers, data sources, and assumptions documented clearly. ## Quick Start To perform market sizing analysis: 1. **Define the market** - Problem, customers, category, geography 2. **Choose methodology** - Bottom-up (preferred) or top-down + triangulation 3. **Gather data** - Industry reports, customer data, competitive intelligence 4. **Calculate TAM** - Apply methodology formula 5. **Narrow to SAM** - Apply product, geographic, segment filters 6. **Estimate SOM** - 2-5% realistic capture rate 7. **Validate** - Cross-check with alternative methods 8. **Document** - Show methodology, sources, assumptions 9. **Present** - Structure for audience (investors, strategy, operations) For detailed step-by-step guidance on each methodology, reference the files in `references/` directory. For complete worked examples, see `examples/` directory.
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

startup-financial-modeling

This skill should be used when the user asks to "create financial

business
⭐1
# Startup Financial Modeling Build comprehensive 3-5 year financial models with revenue projections, cost structures, cash flow analysis, and scenario planning for early-stage startups. ## Overview Financial modeling provides the quantitative foundation for startup strategy, fundraising, and operational planning. Create realistic projections using cohort-based revenue modeling, detailed cost structures, and scenario analysis to support decision-making and investor presentations. ## Core Components ### Revenue Model **Cohort-Based Projections:** Build revenue from customer acquisition and retention by cohort. **Formula:** ``` MRR = Ī£ (Cohort Size Ɨ Retention Rate Ɨ ARPU) ARR = MRR Ɨ 12 ``` **Key Inputs:** - Monthly new customer acquisitions - Customer retention rates by month - Average revenue per user (ARPU) - Pricing and packaging assumptions - Expansion revenue (upsells, cross-sells) ### Cost Structure **Operating Expenses Categories:** 1. **Cost of Goods Sold (COGS)** - Hosting and infrastructure - Payment processing fees - Customer support (variable portion) - Third-party services per customer 2. **Sales & Marketing (S&M)** - Customer acquisition cost (CAC) - Marketing programs and advertising - Sales team compensation - Marketing tools and software 3. **Research & Development (R&D)** - Engineering team compensation - Product management - Design and UX - Development tools and infrastructure 4. **General & Administrative (G&A)** - Executive team - Finance, legal, HR - Office and facilities - Insurance and compliance ### Cash Flow Analysis **Components:** - Beginning cash balance - Cash inflows (revenue, fundraising) - Cash outflows (operating expenses, CapEx) - Ending cash balance - Monthly burn rate - Runway (months of cash remaining) **Formula:** ``` Runway = Current Cash Balance / Monthly Burn Rate Monthly Burn = Monthly Revenue - Monthly Expenses ``` ### Headcount Planning **Role-Based Hiring Plan:** Track headcount by department and role. **Key Metrics:** - Fully-loaded cost per employee - Revenue per employee - Headcount by department (% of total) **Typical Ratios (Early-Stage SaaS):** - Engineering: 40-50% - Sales & Marketing: 25-35% - G&A: 10-15% - Customer Success: 5-10% ## Financial Model Structure ### Three-Scenario Framework **Conservative Scenario (P10):** - Slower customer acquisition - Lower pricing or conversion - Higher churn rates - Extended sales cycles - Used for cash management **Base Scenario (P50):** - Most likely outcomes - Realistic assumptions - Primary planning scenario - Used for board reporting **Optimistic Scenario (P90):** - Faster growth - Better unit economics - Lower churn - Used for upside planning ### Time Horizon **Detailed Projections: 3 Years** - Monthly detail for Year 1 - Monthly detail for Year 2 - Quarterly detail for Year 3 **High-Level Projections: Years 4-5** - Annual projections - Key metrics only - Support long-term planning ## Step-by-Step Process ### Step 1: Define Business Model Clarify revenue model and pricing. **SaaS Model:** - Subscription pricing tiers - Annual vs. monthly contracts - Free trial or freemium approach - Expansion revenue strategy **Marketplace Model:** - GMV projections - Take rate (% of transactions) - Buyer and seller economics - Transaction frequency **Transactional Model:** - Transaction volume - Revenue per transaction - Frequency and seasonality ### Step 2: Build Revenue Projections Use cohort-based methodology for accuracy. **Monthly Customer Acquisition:** Define new customers acquired each month. **Retention Curve:** Model customer retention over time. **Typical SaaS Retention:** - Month 1: 100% - Month 3: 90% - Month 6: 85% - Month 12: 75% - Month 24: 70% **Revenue Calculation:** For each cohort, calculate retained customers Ɨ ARPU for each month. ### Step 3: Model Cost Structure Break down costs by category and behavior. **Fixed vs. Variable:** - Fixed: Salaries, software, rent - Variable: Hosting, payment processing, support **Scaling Assumptions:** - COGS as % of revenue - S&M as % of revenue (CAC payback) - R&D growth rate - G&A as % of total expenses ### Step 4: Create Hiring Plan Model headcount growth by role and department. **Inputs:** - Starting headcount - Hiring velocity by role - Fully-loaded compensation by role - Benefits and taxes (typically 1.3-1.4x salary) **Example:** ``` Engineer: $150K salary Ɨ 1.35 = $202K fully-loaded Sales Rep: $100K OTE Ɨ 1.30 = $130K fully-loaded ``` ### Step 5: Project Cash Flow Calculate monthly cash position and runway. **Monthly Cash Flow:** ``` Beginning Cash + Revenue Collected (consider payment terms) - Operating Expenses Paid - CapEx = Ending Cash ``` **Runway Calculation:** ``` If Ending Cash < 0: Funding Need = Negative Cash Balance Runway = 0 Else: Runway = Ending Cash / Average Monthly Burn ``` ### Step 6: Calculate Key Metrics Track metrics that matter for stage. **Revenue Metrics:** - MRR / ARR - Growth rate (MoM, YoY) - Revenue by segment or cohort **Unit Economics:** - CAC (Customer Acquisition Cost) - LTV (Lifetime Value) - CAC Payback Period - LTV / CAC Ratio **Efficiency Metrics:** - Burn multiple (Net Burn / Net New ARR) - Magic number (Net New ARR / S&M Spend) - Rule of 40 (Growth % + Profit Margin %) **Cash Metrics:** - Monthly burn rate - Runway (months) - Cash efficiency ### Step 7: Scenario Analysis Create three scenarios with different assumptions. **Variable Assumptions:** - Customer acquisition rate (±30%) - Churn rate (±20%) - Average contract value (±15%) - CAC (±25%) **Fixed Assumptions:** - Pricing structure - Core operating expenses - Hiring plan (adjust timing, not roles) ## Business Model Templates ### SaaS Financial Model **Revenue Drivers:** - New MRR (customers Ɨ ARPU) - Expansion MRR (upsells) - Contraction MRR (downgrades) - Churned MRR (lost customers) **Key Ratios:** - Gross margin: 75-85% - S&M as % revenue: 40-60% (early stage) - CAC payback: < 12 months - Net retention: 100-120% **Example Projection:** ``` Year 1: $500K ARR, 50 customers, $100K MRR by Dec Year 2: $2.5M ARR, 200 customers, $208K MRR by Dec Year 3: $8M ARR, 600 customers, $667K MRR by Dec ``` ### Marketplace Financial Model **Revenue Drivers:** - GMV (Gross Merchandise Value) - Take rate (% of GMV) - Net revenue = GMV Ɨ Take rate **Key Ratios:** - Take rate: 10-30% depending on category - CAC for buyers vs. sellers - Contribution margin: 60-70% **Example Projection:** ``` Year 1: $5M GMV, 15% take rate = $750K revenue Year 2: $20M GMV, 15% take rate = $3M revenue Year 3: $60M GMV, 15% take rate = $9M revenue ``` ### E-Commerce Financial Model **Revenue Drivers:** - Traffic (visitors) - Conversion rate - Average order value (AOV) - Purchase frequency **Key Ratios:** - Gross margin: 40-60% - Contribution margin: 20-35% - CAC payback: 3-6 months ### Services / Agency Financial Model **Revenue Drivers:** - Billable hours or projects - Hourly rate or project fee - Utilization rate - Team capacity **Key Ratios:** - Gross margin: 50-70% - Utilization: 70-85% - Revenue per employee ## Fundraising Integration ### Funding Scenario Modeling **Pre-Money Valuation:** Based on metrics and comparables. **Dilution:** ``` Post-Money = Pre-Money + Investment Dilution % = Investment / Post-Money ``` **Use of Funds:** Allocate funding to extend runway and achieve milestones. **Example:** ``` Raise: $5M at $20M pre-money Post-Money: $25M Dilution: 20% Use of Funds: - Product Development: $2M (40%) - Sales & Marketing: $2M (40%) - G&A and Operations: $0.5M (10%) - Working Capital: $0.5M (10%) ``` ### Milestone-Based Planning **Identify Key Milestones:** - Product launch - First $1M ARR - Break-even on CAC - Series A fundraise **Funding Amount:** Ensure runway to achieve next milestone + 6 months buffer. ## Common Pitfalls **Pitfall 1: Overly Optimistic Revenue** - New startups rarely hit aggressive projections - Use conservative customer acquisition assumptions - Model realistic churn rates **Pitfall 2: Underestimating Costs** - Add 20% buffer to expense estimates - Include fully-loaded compensation - Account for software and tools **Pitfall 3: Ignoring Cash Flow Timing** - Revenue ≠ cash (payment terms) - Expenses paid before revenue collected - Model cash conversion carefully **Pitfall 4: Static Headcount** - Hiring takes time (3-6 months to fill roles) - Ramp time for productivity (3-6 months) - Account for attrition (10-15% annually) **Pitfall 5: Not Scenario Planning** - Single scenario is never accurate - Always model conservative case - Plan for what you'll do if base case fails ## Model Validation **Sanity Checks:** - [ ] Revenue growth rate is achievable (3x in Year 2, 2x in Year 3) - [ ] Unit economics are realistic (LTV/CAC > 3, payback < 18 months) - [ ] Burn multiple is reasonable (< 2.0 in Year 2-3) - [ ] Headcount scales with revenue (revenue per employee growing) - [ ] Gross margin is appropriate for business model - [ ] S&M spending aligns with CAC and growth targets **Benchmark Against Peers:** Compare key metrics to similar companies at similar stage. **Investor Feedback:** Share model with advisors or investors for feedback on assumptions. ## Additional Resources ### Reference Files For detailed model structures and advanced techniques: - **`references/model-templates.md`** - Complete financial model templates by business model - **`references/unit-economics.md`** - Deep dive on CAC, LTV, payback, and efficiency metrics - **`references/fundraising-scenarios.md`** - Modeling funding rounds and dilution ### Example Files Working financial models with formulas: - **`examples/saas-financial-model.md`** - Complete 3-year SaaS model with cohort analysis - **`examples/marketplace-model.md`** - Marketplace GMV and take rate projections - **`examples/scenario-analysis.md`** - Three-scenario framework with sensitivities ## Quick Start To create a startup financial model: 1. **Define business model** - Revenue drivers and pricing 2. **Project revenue** - Cohort-based with retention 3. **Model costs** - COGS, S&M, R&D, G&A by month 4. **Plan headcount** - Hiring by role and department 5. **Calculate cash flow** - Revenue - expenses = burn/runway 6. **Compute metrics** - CAC, LTV, burn multiple, runway 7. **Create scenarios** - Conservative, base, optimistic 8. **Validate assumptions** - Sanity check and benchmark 9. **Integrate fundraising** - Model funding rounds and milestones For complete templates and formulas, reference the `references/` and `examples/` files.
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered
šŸ¤–system prompt•7 months ago

team-composition-analysis

This skill should be used when the user asks to "plan team

business
⭐1
# Team Composition Analysis Design optimal team structures, hiring plans, compensation strategies, and equity allocation for early-stage startups from pre-seed through Series A. ## Overview Build the right team at the right time with appropriate compensation and equity. Plan role-by-role hiring aligned with revenue milestones, budget constraints, and market benchmarks. ## Team Structure by Stage ### Pre-Seed (0-$500K ARR) **Team Size: 2-5 people** **Core Roles:** - Founders (2-3): Product, engineering, business - First engineer (if needed) - Contract roles: Design, marketing **Focus:** Build and validate product-market fit ### Seed ($500K-$2M ARR) **Team Size: 5-15 people** **Key Hires:** - Engineering lead + 2-3 engineers - First sales/business development - Product manager - Marketing/growth lead **Focus:** Scale product and prove repeatable sales ### Series A ($2M-$10M ARR) **Team Size: 15-50 people** **Department Build-Out:** - Engineering (40%): 6-20 people - Sales & Marketing (30%): 5-15 people - Customer Success (10%): 2-5 people - G&A (10%): 2-5 people - Product (10%): 2-5 people **Focus:** Scale revenue and build repeatable processes ## Role-by-Role Planning ### Engineering Team **Pre-Seed:** - Founders write code - 0-1 contract developers **Seed:** - Engineering Lead (first $150K-$180K) - 2-3 Full-Stack Engineers ($120K-$150K) - 1 Frontend or Backend Specialist ($130K-$160K) **Series A:** - VP Engineering ($180K-$250K + equity) - 2-3 Senior Engineers ($150K-$180K) - 3-5 Mid-Level Engineers ($120K-$150K) - 1-2 Junior Engineers ($90K-$120K) - 1 DevOps/Infrastructure ($140K-$170K) ### Sales & Marketing **Pre-Seed:** - Founders do sales - Contract marketing help **Seed:** - First Sales Hire / Head of Sales ($120K-$150K + commission) - Marketing/Growth Lead ($100K-$140K) - SDR or BDR (if B2B) ($50K-$70K + commission) **Series A:** - VP Sales ($150K-$200K + commission + equity) - 3-5 Account Executives ($80K-$120K + commission) - 2-3 SDRs/BDRs ($50K-$70K + commission) - Marketing Manager ($90K-$130K) - Content/Demand Gen ($70K-$100K) ### Product Team **Pre-Seed:** - Founder as product lead **Seed:** - First Product Manager ($120K-$150K) - Contract designer **Series A:** - Head of Product ($150K-$180K) - 1-2 Product Managers ($120K-$150K) - Product Designer ($100K-$140K) - UX Researcher (optional) ($90K-$130K) ### Customer Success **Pre-Seed:** - Founders handle support **Seed:** - First CS hire (optional) ($60K-$90K) **Series A:** - CS Manager ($100K-$130K) - 2-4 CS Representatives ($60K-$90K) - Support Engineer (technical) ($80K-$120K) ### G&A (General & Administrative) **Pre-Seed:** - Contractors (accounting, legal) **Seed:** - Operations/Office Manager ($70K-$100K) - Contract CFO **Series A:** - CFO or Finance Lead ($150K-$200K) - Recruiter ($80K-$120K) - Office Manager / EA ($60K-$90K) ## Compensation Strategy ### Base Salary Benchmarks (US, 2024) **Engineering:** - Junior: $90K-$120K - Mid-Level: $120K-$150K - Senior: $150K-$180K - Staff/Principal: $180K-$220K - Engineering Manager: $160K-$200K - VP Engineering: $180K-$250K **Sales:** - SDR/BDR: $50K-$70K base + $50K-$70K commission - Account Executive: $80K-$120K base + $80K-$120K commission - Sales Manager: $120K-$160K base + $80K-$120K commission - VP Sales: $150K-$200K base + $150K-$200K commission **Product:** - Product Manager: $120K-$150K - Senior PM: $150K-$180K - Head of Product: $150K-$180K - VP Product: $180K-$220K **Marketing:** - Marketing Manager: $90K-$130K - Content/Demand Gen: $70K-$100K - Head of Marketing: $130K-$170K - VP Marketing: $150K-$200K **Customer Success:** - CS Representative: $60K-$90K - CS Manager: $100K-$130K - VP Customer Success: $140K-$180K ### Total Compensation Formula ``` Total Comp = Base Salary Ɨ 1.30 (benefits & taxes) + Equity Value ``` **Fully-Loaded Cost:** - Base salary - Payroll taxes (7.65% FICA) - Benefits (health insurance, 401k): $10K-$15K per employee - Other (workspace, equipment, software): $5K-$10K per employee **Rule of Thumb:** Multiply base salary by 1.3-1.4 for fully-loaded cost ### Geographic Adjustments **San Francisco / New York:** +20-30% above benchmarks **Seattle / Boston / Los Angeles:** +10-20% **Austin / Denver / Chicago:** +0-10% **Remote / Other US Cities:** -10-20% **International:** Varies widely by country ## Equity Allocation ### Equity by Role and Stage **Founders:** - First founder: 40-60% - Second founder: 20-40% - Third founder: 10-20% - Vesting: 4 years with 1-year cliff **Early Employees (Pre-Seed):** - First engineer: 0.5-2.0% - First 5 employees: 0.25-1.0% each **Seed Stage Hires:** - VP/Head level: 0.5-1.5% - Senior IC: 0.1-0.5% - Mid-level: 0.05-0.25% - Junior: 0.01-0.1% **Series A Hires:** - C-level (CTO, CFO): 1.0-3.0% - VP level: 0.3-1.0% - Director level: 0.1-0.5% - Senior IC: 0.05-0.2% - Mid-level: 0.01-0.1% - Junior: 0.005-0.05% ### Equity Pool Sizing **Option Pool by Round:** - Pre-Seed: 10-15% reserved - Seed: 10-15% top-up - Series A: 10-15% top-up - Series B+: 5-10% per round **Pre-Funding Dilution:** Investors often require option pool creation before investment, diluting founders. **Example:** ``` Pre-money: $10M Investors want 15% option pool post-money Calculation: Post-money: $15M ($10M + $5M investment) Option pool: $2.25M (15% Ɨ $15M) Founders diluted by pool creation before new money ``` ## Organizational Design ### Reporting Structure **Pre-Seed:** ``` Founders (flat structure) ā”œā”€ā”€ Contractors └── First hires (report to founders) ``` **Seed:** ``` CEO ā”œā”€ā”€ Engineering Lead (2-4 engineers) ā”œā”€ā”€ Sales/Growth Lead (1-2 reps) ā”œā”€ā”€ Product Manager └── Operations ``` **Series A:** ``` CEO ā”œā”€ā”€ CTO / VP Engineering (6-20 people) │ ā”œā”€ā”€ Engineering Manager(s) │ └── Individual Contributors ā”œā”€ā”€ VP Sales (5-15 people) │ ā”œā”€ā”€ Sales Manager │ ā”œā”€ā”€ Account Executives │ └── SDRs ā”œā”€ā”€ Head of Product (2-5 people) │ ā”œā”€ā”€ Product Managers │ └── Designers ā”œā”€ā”€ Head of Customer Success (2-5 people) └── CFO / Finance Lead (2-5 people) ā”œā”€ā”€ Recruiter └── Operations ``` ### Span of Control **Manager Ratios:** - First-line managers: 4-8 direct reports - Directors: 3-5 direct reports (managers) - VPs: 3-5 direct reports (directors) - CEO: 5-8 direct reports (executive team) ## Full-Time vs. Contract ### Use Full-Time for: - Core product development - Sales (revenue-generating roles) - Mission-critical operations - Institutional knowledge roles ### Use Contractors for: - Specialized short-term needs (legal, accounting) - Variable workload (design, marketing campaigns) - Skills outside core competency - Testing role before FTE hire - Geographic expansion before permanent presence ### Cost Comparison **Full-Time:** - Lower hourly cost - Benefits and overhead - Long-term commitment - Cultural fit matters **Contract:** - Higher hourly rate ($75-$200/hour vs. $40-$100/hour FTE equivalent) - No benefits or overhead - Flexible engagement - Easier to scale up/down ## Hiring Velocity ### Realistic Timeline **Role Opening to Hire:** - Junior: 6-8 weeks - Mid-Level: 8-12 weeks - Senior: 12-16 weeks - Executive: 16-24 weeks **Time to Productivity:** - Junior: 4-6 months - Mid-Level: 2-4 months - Senior: 1-3 months - Executive: 3-6 months ### Planning Buffer Always add 2-3 months buffer to hiring plans. **Example:** If need engineer by July 1: - Start recruiting: April 1 (12 weeks) - Productivity: September 1 (2 months ramp) ## Budget Planning ### Compensation as % of Revenue **Early Stage (Seed):** - Total comp: 120-150% of revenue (burning cash to grow) - Engineering: 50-60% - Sales: 30-40% - Other: 20-30% **Growth Stage (Series A):** - Total comp: 70-100% of revenue - Engineering: 35-45% - Sales: 25-35% - Other: 20-30% ### Headcount Budget Formula ``` Total Comp Budget = Ī£ (Role Count Ɨ Fully-Loaded Cost Ɨ % of Year) Example: 3 Engineers Ɨ $202K Ɨ 100% = $606K 2 AEs Ɨ $230K Ɨ 75% (mid-year start) = $345K 1 PM Ɨ $162K Ɨ 100% = $162K Total: $1.1M ``` ## Additional Resources ### Reference Files - **`references/compensation-benchmarks.md`** - Detailed salary data by role, level, and location - **`references/equity-calculator.md`** - Equity sizing formulas and dilution scenarios ### Example Files - **`examples/seed-stage-hiring-plan.md`** - Complete hiring plan for seed-stage SaaS company - **`examples/org-chart-evolution.md`** - Organizational design from 5 to 50 people ## Quick Start To plan team composition: 1. **Identify stage** - Pre-seed, seed, or Series A 2. **Define roles** - What functions are needed now 3. **Prioritize hires** - Critical path for business goals 4. **Set compensation** - Base salary + equity by level 5. **Plan timeline** - Account for recruiting and ramp time 6. **Calculate budget** - Fully-loaded cost Ɨ headcount 7. **Design org chart** - Reporting structure and span of control 8. **Allocate equity** - Fair allocation that preserves pool For detailed compensation benchmarks and hiring plan templates, see `references/` and `examples/`.
šŸ‘0
šŸ‘ļø0
šŸ¤– Auto-discovered