Why Sample Size Matters
Educational Resource: This guide explains statistical sample size concepts used in quantitative analysis. These concepts help evaluate historical pattern reliability but do not guarantee future performance. All investments involve risk.
The Sample Size Question
Every driver shows you an "n" value - the number of times we've tested that pattern. But why does this number matter so much?
Simple answer: Anyone can flip a coin 3 times and get 3 heads. That doesn't mean the coin is rigged. But if you flip it 1,000 times and get 750 heads, something is definitely unusual.
Sample size (n) tells you whether a pattern is luck or real.
What is Sample Size?
Sample size is the number of historical occurrences we've tested for a specific pattern.
Example:
insider_buying_20d → forward_return_30d
r=0.68, n=176
Translation: We found 176 instances in historical data where insiders bought heavily over 20 days. We measured what happened to the stock 30 days later in all 176 cases. The correlation was 0.68.
The larger the n, the more confident we can be that r=0.68 is a real relationship, not random chance.
Why Small Samples Are Dangerous
The Coin Flip Example
Imagine testing a "pattern" by flipping a coin:
Test 1: n=3 flips
- Results: Heads, Heads, Heads
- Success rate: 100%
- Conclusion: The coin always lands heads!
Test 2: n=100 flips
- Results: 52 heads, 48 tails
- Success rate: 52%
- Conclusion: The coin is basically random
With n=3, we thought we had a "can't lose" pattern. With n=100, we learned the truth: it's a coin.
The same thing happens with trading patterns. Small samples make random noise look like reliable signals.
Real Trading Example
Pattern: "Insider buying by CEOs → stock rises"
Test 1: Small Sample (n=12)
- Tested on 12 companies
- 10 stocks rose (83% success rate!)
- Correlation: r=0.76
- Conclusion: Amazing pattern!
Test 2: Large Sample (n=200)
- Tested on 200 companies
- 110 stocks rose (55% success rate)
- Correlation: r=0.42
- Conclusion: Barely better than coin flip
What happened? With n=12, we got lucky. With n=200, we found the true relationship: modest and inconsistent.
This is why we require n ≥ 30 minimum - to filter out lucky noise.
The Relationship: Sample Size vs Confidence
As sample size increases, our confidence in the pattern increases:
| Sample Size | Confidence Level | Interpretation |
|---|---|---|
| n < 30 | Very Low | Could easily be luck - we don't show these |
| n = 30-50 | Low-Moderate | Emerging pattern - use with caution |
| n = 50-100 | Moderate | Decent evidence - verify with other signals |
| n = 100-200 | High | Strong evidence - reliable pattern |
| n ≥ 200 | Very High | Robust pattern - well-tested across many scenarios |
Think of it like a weather forecast:
- Forecast based on 5 days of data: Not reliable
- Forecast based on 50 years of data: Much more reliable
Visual: Confidence Progression
As you gather more evidence, confidence in the pattern grows exponentially:
Think of it like building a legal case: 3 witnesses might be coincidence, but 200 witnesses telling the same story? That's proof.
Sample Size and Correlation Strength Together
Sample size and correlation strength (r-value) work together to determine pattern quality:
Scenario 1: Strong Correlation, Small Sample
r = 0.85, n = 25
Interpretation:
- Very strong correlation... but
- Small sample - could be lucky
- Risk: Pattern might not hold up with more data
- Action: Use caution, require additional confirmation
Scenario 2: Moderate Correlation, Large Sample
r = 0.52, n = 250
Interpretation:
- Moderate correlation
- Large sample - highly confident it's real
- Risk: Returns will be modest and variable
- Action: Reliable but not explosive - good for consistent gains
Scenario 3: Strong Correlation, Large Sample
r = 0.74, n = 180
Interpretation:
- Strong correlation
- Large sample - very confident it's real
- Risk: Lower risk of the pattern being luck
- Action: High-quality setup - prioritize these
Scenario 4: Weak Correlation, Any Sample
r = 0.22, n = 500
Interpretation:
- Weak correlation even with massive sample
- Pattern exists but barely
- Risk: Not worth trading
- Action: Filter out - we don't show these (r < 0.3)
Quick Reference: Pattern Quality Scenarios
| Scenario | r-value | Sample (n) | Verdict | Action | Risk Level |
|---|---|---|---|---|---|
| 🚩 Lucky Streak | 0.85 | 25 | Possibly luck | ⚠️ Use extreme caution - require multiple confirmations | High |
| ✅ Reliable Moderate | 0.52 | 250 | Real but modest | Trade with confidence - expect consistent, moderate returns | Low-Medium |
| 🌟 Elite Setup | 0.74 | 180 | Highly confident | Prioritize these - strongest patterns in the system | Low |
| ❌ Weak Pattern | 0.22 | 500 | Not actionable | Filter out - not worth trading even with large sample | N/A |
Key insight: A moderate correlation with a large sample (Scenario 2) is often more reliable than a strong correlation with a small sample (Scenario 1). Sample size trumps correlation strength when assessing reliability.
How We Use Sample Size Standards
Minimum Threshold (n ≥ 30)
We won't show you a pattern unless we've tested it at least 30 times. This filters out:
- Lucky streaks
- Rare events that might not repeat
- Overfitted patterns from data mining
Why 30? It's the statistical minimum for the "Law of Large Numbers" to start working. Below 30, randomness dominates.
Preferred Range (n ≥ 100)
For high-conviction trades, look for n ≥ 100. At this sample size:
- Randomness is largely eliminated
- Pattern has been tested across multiple market conditions
- Statistical significance is much stronger
- Pattern likely works in different sectors and time periods
Elite Patterns (n ≥ 200)
When you see n ≥ 200, you're looking at battle-tested patterns that have:
- Survived bull markets and bear markets
- Worked across many different stocks
- Held up over many years
- Very low chance of being luck
Sample Size and Time Periods
Sample size also reflects how much historical time we've covered:
Example Analysis
Pattern: Insider buying before earnings Sample size: n=150 Time period: 2015-2024 (10 years)
What this means:
- We found 150 instances of this pattern over 10 years
- That's ~15 instances per year
- Pattern has worked in different market cycles:
- 2015-2019: Bull market
- 2020: COVID crash and recovery
- 2021-2022: Volatility and bear market
- 2023-2024: Recovery
Conclusion: Pattern survived multiple regimes. More robust than a pattern that only worked in one type of market.
When Small Samples Are OK
There are a few exceptions where smaller samples (n=30-50) might still be valuable:
1. Sector-Specific Patterns
In niche industries (biotech, crypto, etc.), events are rarer. A pattern with n=40 in biotech might be well-tested if that industry only has ~50 companies.
How to verify: Check if n represents a significant % of available opportunities in that sector.
2. Recent Regime Changes
If market structure changed recently (new regulations, algorithm trading dominance), older data might not apply. A pattern with n=50 from recent years might be more relevant than n=200 spanning 20 years.
How to verify: Check the time period tested. Recent patterns (2020-2024) might be more applicable than patterns from 2000-2020.
3. Very Strong Correlations
Sometimes the relationship is so strong (r ≥ 0.85) that even n=40 provides confidence.
How to verify: Check if r-value is exceptional (≥0.80) AND p-value is very low (p < 0.001).
Still prefer larger samples when possible.
Statistical Significance and Sample Size
Sample size directly affects statistical significance (p-value):
Same Correlation, Different Samples
Pattern A:
r = 0.50, n = 30, p = 0.08 (not significant)
Pattern B:
r = 0.50, n = 200, p < 0.001 (highly significant)
Same correlation strength, but Pattern B has a much larger sample, making it statistically significant.
Why? With n=30, r=0.50 could happen by luck. With n=200, r=0.50 is almost certainly real.
Learn more: Statistical Significance
Practical Application: Evaluating Sample Size
Step 1: Check Minimum Threshold
Before anything else, verify n ≥ 30. If it's below 30, the pattern didn't pass our filters (you won't see it).
Step 2: Categorize Confidence
Use our confidence levels:
- n < 50: Low confidence (use with caution)
- n = 50-100: Moderate confidence (good for confirmation)
- n = 100-200: High confidence (reliable patterns)
- n ≥ 200: Very high confidence (elite patterns)
Step 3: Combine with R-Value
Evaluate quality using both metrics:
High Quality: r ≥ 0.6 AND n ≥ 100 Moderate Quality: r ≥ 0.5 AND n ≥ 50 Use Caution: r < 0.5 OR n < 50
Step 4: Check Time Period
If available, verify the pattern was tested across multiple years and market conditions.
Prefer: Patterns tested 2015-2024 (10 years, multiple regimes) Caution: Patterns tested only 2020-2021 (might be COVID-specific)
Real Example: Comparing Two Opportunities
Opportunity A: High Score, Small Sample
Top Driver:
insider_cluster_buying → forward_return_20d
r = 0.82, n = 38, contribution = 45%
Analysis:
- Strong correlation (0.82)
- Small sample (38) - could be sector-specific or lucky
- High contribution (45%) - dominates the score
- Risk: Pattern not well-tested, might be overfitted
Decision: Proceed with caution. Require additional confirming signals before trading.
Opportunity B: Moderate Score, Large Sample
Top Driver:
insider_buying_consistent → forward_return_30d
r = 0.58, n = 187, contribution = 28%
Analysis:
- Moderate correlation (0.58)
- Large sample (187) - well-tested and robust
- Moderate contribution (28%)
- Risk: Lower expected returns, but higher reliability
Decision: More consistent pattern. Better for risk-averse traders.
Which is Better?
It depends on your strategy:
Aggressive traders: Might prefer Opportunity A (higher potential, higher risk) Conservative traders: Might prefer Opportunity B (lower potential, lower risk)
Best practice: Diversify - take some of each if both pass your other criteria.
Common Questions
Q: Is n=100 always better than n=50?
A: Generally yes, but context matters. n=50 from the last 2 years might be more relevant than n=100 spanning 20 years if market structure changed recently.
Q: Can sample size be too large?
A: Rarely a problem, but if n=1000+ spans 30+ years, some of that data might be outdated. Prefer recent, regime-appropriate samples over ancient data.
Q: Why don't you show patterns with n < 30?
A: Because they're unreliable. We tested thousands of patterns with n=10-30 and found they don't hold up when tested on fresh data. Setting n ≥ 30 eliminates most false positives.
Q: How can I see the sample size for a driver?
A: Click into the opportunity's detail panel. Each driver shows: r-value, sample size (n), and contribution percentage.
Q: Does sample size affect the opportunity score?
A: Yes, indirectly. Patterns with larger samples tend to have more stable correlations, which increases their contribution to the overall score. We also weight high-sample patterns more heavily in our algorithms.
Next Steps
Ready to understand the statistical tests we use to validate patterns? Learn about Statistical Significance and how p-values work.
Want to see how we prevent false discoveries? Check out Multiple Testing Correction.