Part 2 of the Meta Data Lineage Series
Important Note: This article is based on my understanding after reading the Meta Engineering blog post and several related articles. I’m trying to demystify and explain these concepts in an accessible way. If you want to understand exactly what Meta built, please refer to the original article linked in the Further Reading section.
The challenge recap: Meta needed to trace data flows across millions of assets, billions of lines of code, and thousands of engineers — all updating in real-time.
This article reveals the technical architecture, clever techniques, and hard-won lessons from building data lineage at unprecedented scale.
Previous: ← Part 1 - The Challenge
Introduction: The Architecture
Remember the problem from Part 1: When a user shares a family photo with location data, how do you track where that data goes across Meta’s entire infrastructure?
Meta’s solution: A two-stage data lineage system
Stage 1: Collect Data Flow Signals
├─ Static code analysis
├─ Runtime instrumentation (Privacy Probes)
├─ SQL query analysis
├─ Config parsing
└─ Combines signals from all sources
Stage 2: Identify Relevant Flows
├─ Start from source assets (e.g., religion data)
├─ Traverse lineage graph
├─ Filter relevant flows
├─ Apply privacy controls
└─ Continuous verification
Result: End-to-end data flow visibility
Let’s walk through each part using the photo location data example.
Stage 1: Collecting Data Flow Signals
Overview: Three Different Approaches
Meta uses different techniques for different system types:
(Web, Backend Services)"] BATCH["Batch Processing
(Data Warehouse)"] AI["AI Systems
(ML Models)"] STATIC["Static Code Analysis"] PROBES["Privacy Probes
(Runtime)"] SQL["SQL Analysis"] CONFIG["Config Parsing"] LINEAGE["Lineage Graph"] SYSTEMS --> WEB SYSTEMS --> BATCH SYSTEMS --> AI WEB --> STATIC WEB --> PROBES BATCH --> SQL AI --> CONFIG STATIC --> LINEAGE PROBES --> LINEAGE SQL --> LINEAGE CONFIG --> LINEAGE style PROBES fill:#fef3c7,stroke:#f59e0b style LINEAGE fill:#dbeafe,stroke:#2563eb
Why three approaches? Different systems require different techniques:
| System Type | Example | Best Approach | Why |
|---|---|---|---|
| Web/APIs | Hack, C++, Python | Static + Runtime | Complex transformations |
| Data Warehouse | SQL queries | SQL Analysis | Declarative, easy to parse |
| AI Systems | Model training | Config Parsing | Configs define relationships |
Collecting Signals: Web Systems
Step 1: User Shares Photo with Location
// Photo upload endpoint (Hack code)
function handlePhotoUpload(user_id: int, photo_data: array): void {
// User shares photo with location
$location = $photo_data['location']; // "Home, 123 Main St"
// Store in database
$this->db->insert('photos', [
'user_id' => $user_id,
'photo_url' => $photo_data['url'],
'location' => $location,
]);
// Log the upload
$this->logger->log('photo_uploaded', [
'user_id' => $user_id,
'location' => $location,
]);
// Call nearby friends service
$this->api->call('/friends/nearby', [
'location' => $location,
]);
}
Question: How do we track where $location goes?
Technique 1: Static Code Analysis
How it works:
Static analyzer simulates code execution:
1. Parse code into Abstract Syntax Tree (AST)
2. Track data flow through variables
3. Identify "sources" (inputs) and "sinks" (outputs)
4. Build flow graph
For our example:
$location (source)
├─> $this->db->insert (sink: database)
├─> $this->logger->log (sink: logs)
└─> $this->api->call (sink: API)
Result: 3 potential data flows detected
The power:
Advantages:
✅ Analyzes without running code
✅ Covers all code paths (even unexecuted)
✅ Fast (can analyze millions of lines)
✅ Finds hidden flows
Limitations:
❌ False positives (code that never actually runs)
❌ Can't see runtime transformations
❌ Struggles with dynamic code
Example false positive:
if (DEBUG_MODE) { // Never true in production
$this->log('debug', $location);
}
Static analysis says: "Location logged here!"
Reality: This code never runs
Need runtime signals to verify!
Technique 2: Privacy Probes (Runtime Instrumentation)
The breakthrough innovation:
Privacy Probes instrument Meta’s core frameworks to capture actual data flows at runtime.
How it works:
// Meta's logging framework (instrumented)
class Logger {
public function log(string $event, array $data): void {
// NEW: Privacy Probe captures this
PrivacyProbe::captureSource($data); // Capture before logging
// Original logging code
$this->writeToLog($event, $data);
// NEW: Privacy Probe captures this
PrivacyProbe::captureSink($event, $data); // Capture after logging
}
}
The magic: Payload matching
Privacy Probes compare source and sink payloads to verify data flows:
Request execution:
1. User submits: location = "Home, 123 Main St"
2. Privacy Probe captures SOURCE:
- Location: handlePhotoUpload()
- Payload: {location: "Home, 123 Main St"}
- Timestamp: 10:30:00.123
3. Code executes...
4. Privacy Probe captures SINK #1:
- Location: db.insert('photos')
- Payload: {user_id: 123, location: "Home, 123 Main St"}
- Timestamp: 10:30:00.145
5. Privacy Probe captures SINK #2:
- Location: logger.log('photo_uploaded')
- Payload: {user_id: 123, location: "Home, 123 Main St"}
- Timestamp: 10:30:00.156
6. Compare payloads:
SOURCE: "Home, 123 Main St"
SINK #1: "Home, 123 Main St" ← EXACT_MATCH ✅
SINK #2: "Home, 123 Main St" ← EXACT_MATCH ✅
7. Emit lineage signals:
- Flow confirmed: location → photos table
- Flow confirmed: location → photo_uploaded logs
Handling transformations:
Case 1: Exact copy
Source: "Home, 123 Main St"
Sink: "Home, 123 Main St"
Result: EXACT_MATCH (high confidence) ✅
Case 2: Contained in larger structure
Source: "Home, 123 Main St"
Sink: {metadata: {location: "Home, 123 Main St", lat: 40.7, lon: -74.0}}
Result: CONTAINS (high confidence) ✅
Case 3: Transformed
Source: {locations: ["Home", "Work", "School"]}
Sink: {location_count: 3}
Result: NO_MATCH (low confidence) ⚠️
Match-set vs Full-set:
Match-set (high confidence):
├─ EXACT_MATCH: Source equals sink
├─ CONTAINS: Sink contains source as substring
└─ Used for automatic lineage
Full-set (lower confidence):
├─ All source-sink pairs in a request
├─ Includes transformed data (NO_MATCH)
├─ Requires human review
└─ Used for discovering hidden transformations
The Power of Runtime + Static Analysis
Combined approach:
Static Analysis finds:
- All possible code paths (100 potential flows)
- Including paths never executed
- High recall, lower precision
Privacy Probes verify:
- Actual flows that happen (15 real flows)
- With exact payloads
- High precision, lower recall (only sampled)
Combined:
✅ High recall (static finds everything)
✅ High precision (runtime verifies)
✅ Best of both worlds
Real example:
function processPhoto(data: array): void {
if (should_log_location()) { // Complex business logic
log('location', data['location']);
}
if (should_send_to_ads()) { // More complex logic
api->call('/ad_targeting', data);
}
}
Static analysis says:
"Location might flow to logs AND ad targeting"
Privacy Probes observe (over 1 week):
- Logs: Observed 10,000 times ✅ Definitely happens
- Ad targeting: Observed 0 times ❌ Likely gated/blocked
Conclusion:
- Log flow: REAL (include in lineage)
- Ad targeting flow: FALSE POSITIVE (exclude)
Instrumentation Points
Where Privacy Probes are embedded:
Meta's core frameworks:
├─ Database layers (MySQL, RocksDB)
│ └─ Captures: Table reads/writes
│
├─ Logging frameworks
│ └─ Captures: All log writes
│
├─ API clients
│ └─ Captures: API calls
│
├─ Caching layers (Memcache, TAO)
│ └─ Captures: Cache reads/writes
│
├─ Message queues
│ └─ Captures: Queue reads/writes
│
└─ Service mesh
└─ Captures: Service-to-service calls
Coverage: 90%+ of data flows
Collecting Signals: Data Warehouse (SQL Systems)
The Challenge
-- Data warehouse jobs process billions of rows
-- How do we track lineage through SQL?
-- Example: Location analytics job
INSERT INTO location_analytics_tbl
SELECT
user_id as target_user_id,
location as user_location,
photo_count,
CASE
WHEN location LIKE '%Home%' THEN 'residential'
ELSE 'public'
END as location_type
FROM photo_metadata_log
WHERE date >= '2025-01-01'
AND user_id IS NOT NULL;
Questions:
- Which columns flow where?
- What transformations happen?
- Where does
locationend up?
Technique: SQL Query Analysis
How it works:
1. SQL queries are logged by compute engines:
- Presto
- Spark
- Hive
- Others
2. Static SQL analyzer parses queries:
- Extract input tables
- Extract output tables
- Map column lineage
3. Build lineage graph:
Input: dating_profiles_log.religion
Output: safety_training_tbl.target_religion
Transformation: Direct copy + rename
SQL analyzer example:
-- Original query
SELECT location as user_location
FROM photo_metadata_log
-- Analyzer output
{
"input_table": "photo_metadata_log",
"input_columns": ["location"],
"output_table": "location_analytics_tbl",
"output_columns": ["user_location"],
"column_lineage": [
{
"source": "photo_metadata_log.location",
"target": "location_analytics_tbl.user_location",
"transformation": "RENAME"
}
],
"confidence": "HIGH"
}
Column-level lineage:
Granular tracking:
Table-level:
photo_metadata_log → location_analytics_tbl
(Not very useful: what about other columns?)
Column-level:
photo_metadata_log.location → location_analytics_tbl.user_location
photo_metadata_log.user_id → location_analytics_tbl.target_user_id
(Much more useful for privacy!)
Handling Complex SQL
Challenge: Complex transformations
-- Complex transformation
SELECT
user_id,
-- Direct copy
location,
-- Aggregation
COUNT(*) as photo_count,
-- CASE statement
CASE
WHEN location LIKE '%Home%' THEN 1
ELSE 0
END as is_home,
-- Concatenation
CONCAT(location, '_', city) as full_location,
-- Subquery
(SELECT COUNT(*) FROM friends WHERE friends.user_id = p.user_id) as friend_count
FROM photo_metadata_log p;
SQL analyzer handles:
For each output column, track dependencies:
location → location (direct copy) ✅
location → is_home (derived) ✅
location → full_location (concatenation) ✅
Photo count:
- Depends on: COUNT(*) aggregate
- Does NOT contain location data ✅
Friend_count:
- Depends on: subquery
- Need to analyze subquery separately ✅
Connecting Reads and Writes
The challenge:
Job logs might be incomplete:
Log 1: "Read from photo_metadata_log"
[No write logged]
Log 2: "Write to location_analytics_tbl"
[No read logged]
Question: Are these the same job?
Solution: Contextual matching
Use execution context to connect reads/writes:
Context matching:
├─ Job ID (same Spark job)
├─ Trace ID (same execution trace)
├─ Timestamp (within same time window)
├─ User ID (same service account)
└─ Execution environment
Example:
Log 1:
job_id: spark_job_12345
action: READ photo_metadata_log
timestamp: 10:30:00
Log 2:
job_id: spark_job_12345 ← Same job!
action: WRITE location_analytics_tbl
timestamp: 10:30:15
Conclusion: These are connected!
Flow: photo_metadata_log → location_analytics_tbl ✅
Collecting Signals: AI Systems
The Challenge
# Model training config
training_config = {
"model_name": "nearby_friends_model",
"input_dataset": "asset://hive.table/photo_location_features",
"features": [
"USER_AGE",
"USER_ACTIVITY_LEVEL",
"USER_HOME_LOCATION", # ← Home location feature!
],
"output_model": "asset://ai.model/nearby_friends_model",
"training_params": {...}
}
Questions:
- Which data does this model use?
- Does it use sensitive data?
- Where are the model inferences used?
Technique: Config Parsing + Runtime Instrumentation
Config parsing:
Parse training configs to extract relationships:
Input:
- Dataset: photo_location_features
- Features: USER_HOME_LOCATION
Output:
- Model: nearby_friends_model
Lineage:
photo_location_features
→ USER_HOME_LOCATION
→ nearby_friends_model ✅
Full AI lineage chain:
photo_location_features"] FEATURE["Feature
USER_HOME_LOCATION"] MODEL["Model
nearby_friends_model"] INFERENCE["Inference Service
friend_suggestions"] PRODUCT["Product
Facebook App"] DATA --> FEATURE FEATURE --> MODEL MODEL --> INFERENCE INFERENCE --> PRODUCT style FEATURE fill:#fef3c7,stroke:#f59e0b style MODEL fill:#dbeafe,stroke:#2563eb
Capturing at multiple levels:
AI pipeline stages:
1. Data loading (DPP framework)
- Instrumented to capture: dataset → features
2. Model training (FBLearner Flow)
- Instrumented to capture: features → model
3. Model registration
- Config parsing: model metadata
4. Inference service (Backend service)
- Privacy Probes: model → API responses
5. Product integration
- Privacy Probes: API → user-facing features
Result: End-to-end AI lineage ✅
Real Example: AI Lineage
Starting point: Location data in photo_location_features
Lineage trace:
1. Data loading:
photo_location_features.home_location
→ DPP loads into feature store
2. Feature engineering:
feature_store.home_location
→ USER_HOME_LOCATION (computed feature)
3. Model training:
USER_HOME_LOCATION
→ nearby_friends_model (training input)
4. Model deployment:
nearby_friends_model
→ inference_service (model serving)
5. Product:
inference_service
→ Facebook friend suggestions
6. Privacy check:
Is home location used for ads?
- Check inference_service callers
- If ad_targeting calls it: VIOLATION ❌
- If only friend_suggestions: COMPLIANT ✅
Stage 2: Identifying Relevant Data Flows
The Lineage Graph
After collecting signals, Meta has a massive graph:
Scale:
├─ Nodes: 10,000,000+ assets
├─ Edges: 100,000,000+ data flows
├─ Updates: Real-time (1,000+ per second)
└─ Storage: Graph database
Challenge: How do you query this efficiently?
The Iterative Discovery Process
Problem: Given a source (photo location data), find all relevant downstream flows.
Naive approach (doesn’t work):
Start at: dating_profiles.religion
Find: All downstream assets
Result:
- 1,000,000+ downstream assets
- 99.9% are false positives
- Takes hours to compute
- Unusable
Why? The graph is too large and interconnected!
Meta’s solution: Iterative filtering
3-step cycle:
1. Discover flows:
- Start from source
- Traverse graph
- Stop at low-confidence edges
2. Human review:
- Engineer excludes false positives
- Engineer confirms true positives
- Apply privacy controls
3. Repeat:
- Use confirmed nodes as new sources
- Continue traversal
- Until no new nodes
Result: Only relevant flows, manageable size
Example: Finding Photo Location Data Flows
Iteration 1:
Source: photos.location
Discover (automatic):
├─ photo_metadata_log (high confidence) ✅
├─ photo_upload_events (high confidence) ✅
├─ nearby_friends_temp_table (low confidence) ⚠️
└─ generic_analytics_table (low confidence) ⚠️
Human review:
├─ photo_metadata_log: INCLUDE ✅
│ └─ Action: Apply privacy controls
│
├─ photo_upload_events: INCLUDE ✅
│ └─ Action: Apply privacy controls
│
├─ nearby_friends_temp_table: EXCLUDE ❌
│ └─ Reason: Only uses approximate location
│
└─ generic_analytics_table: EXCLUDE ❌
└─ Reason: Aggregated, no precise locations
Confirmed nodes: 2 (photo_metadata_log, photo_upload_events)
Iteration 2:
New sources: photo_metadata_log, photo_upload_events
Discover from photo_metadata_log:
├─ location_analytics_tbl (high confidence) ✅
├─ data_warehouse_agg_daily (low confidence) ⚠️
└─ export_table_xyz (low confidence) ⚠️
Human review:
├─ location_analytics_tbl: INCLUDE ✅
│ └─ Action: Apply privacy controls
│
└─ Others: EXCLUDE ❌
Confirmed nodes: +1 (location_analytics_tbl)
Iteration 3:
New sources: location_analytics_tbl
Discover from location_analytics_tbl:
├─ ai_location_feature_store (high confidence) ✅
└─ No other high-confidence flows
Human review:
└─ ai_location_feature_store: INCLUDE ✅
Confirmed nodes: +1 (ai_location_feature_store)
Iteration 4:
New sources: ai_location_feature_store
Discover:
├─ nearby_friends_model (high confidence) ✅
└─ ad_targeting_model (high confidence) ⚠️
Human review:
├─ nearby_friends_model: INCLUDE ✅
│ └─ Final destination: Friend suggestions only
│ └─ Privacy requirement: SATISFIED ✅
│
└─ ad_targeting_model: VIOLATION ❌
└─ Home location used for ads!
└─ Action: BLOCK and alert teams
Done! Privacy violation detected and prevented.
The Power of Cascading Exclusions
Key insight:
When you exclude a node early, all its downstream is automatically excluded — saving massive review effort.
Example:
If generic_analytics_table is excluded in Iteration 1:
├─ Excludes itself: 1 asset
├─ Excludes 500 downstream tables
├─ Excludes 10,000 downstream views
└─ Excludes 100,000 total assets
Total review effort saved: 99%+ ✅
By excluding one false positive early,
you avoid reviewing 100,000 downstream assets!
The Tool: Policy Zone Manager (PZM)
Developer Experience
Before PZM (manual):
Engineer task: "Protect photo location data"
Week 1-4: Find where location is used
- Read code manually
- Interview teams
- Build spreadsheet
Week 5-8: Identify downstream flows
- Trace each flow manually
- Check transformations
- More spreadsheets
Week 9-12: Apply privacy controls
- Modify code in 100+ places
- Test each change
- Hope nothing breaks
Total: 3 months, high error rate ❌
After PZM (with lineage):
Engineer task: "Protect photo location data"
Hour 1: Query lineage
- Open PZM tool
- Query: "Find all flows from photos.location"
- Results: Visual graph in 30 seconds ✅
Hour 2-4: Review flows
- Interactive UI shows all flows
- Mark relevant: Include/Exclude
- Tool applies controls automatically
Day 1: Deploy controls
- PZM generates code changes
- Review and deploy
- Continuous monitoring enabled
Total: 1 day, low error rate ✅
Time saved: 99%
PZM Features
Visual lineage graph:
Interactive UI:
├─ Visual graph of data flows
├─ Click nodes to expand/collapse
├─ Color coding:
│ ├─ Green: Protected
│ ├─ Yellow: Needs review
│ └─ Red: Violation detected
├─ Filter by:
│ ├─ Confidence level
│ ├─ Data type
│ └─ System
└─ Export to code changes
Bulk operations:
Instead of reviewing one flow at a time:
Bulk exclude:
- Select 100 false positive nodes
- Click "Exclude all + downstream"
- Saves weeks of work ✅
Bulk include:
- Select all high-confidence Dating assets
- Click "Apply privacy controls"
- Automatically generates code ✅
Continuous monitoring:
After initial setup:
PZM monitors continuously:
├─ New code deployed
├─ New data flows detected
├─ Check against policies
├─ Alert if violation:
│ └─ "New flow detected: home_location → ad_targeting"
│ └─ "BLOCK: Violates purpose limitation"
└─ Automatic enforcement ✅
The Technology Stack
Architecture Overview
Servers"] BATCH["Data
Warehouse"] AI["AI
Systems"] COLLECTORS["Collectors"] STATIC["Static
Analyzer"] PROBES["Privacy
Probes"] SQL["SQL
Analyzer"] PIPELINE["Processing Pipeline"] DEDUP["Deduplication"] SCORING["Confidence
Scoring"] MERGE["Signal
Merging"] GRAPH["Lineage Graph
(Graph DB)"] TOOLS["Tools"] PZM["Policy Zone
Manager"] API["Query
API"] MONITOR["Continuous
Monitoring"] SOURCES --> WEB SOURCES --> BATCH SOURCES --> AI WEB --> STATIC WEB --> PROBES BATCH --> SQL AI --> STATIC STATIC --> COLLECTORS PROBES --> COLLECTORS SQL --> COLLECTORS COLLECTORS --> PIPELINE PIPELINE --> DEDUP DEDUP --> SCORING SCORING --> MERGE MERGE --> GRAPH GRAPH --> PZM GRAPH --> API GRAPH --> MONITOR style PROBES fill:#fef3c7,stroke:#f59e0b style GRAPH fill:#dbeafe,stroke:#2563eb
Key Components
1. Privacy Probes (Runtime)
Implementation:
├─ Language: C++ (performance-critical)
├─ Sampling: 1% of requests (configurable)
├─ Overhead: <1% latency impact
├─ Storage: In-memory buffers
└─ Processing: Asynchronous
Scale:
├─ Requests sampled: 10M+ per second
├─ Payloads captured: 100M+ per second
├─ Comparisons: 1B+ per second
└─ Signals emitted: 10M+ per second
2. Static Analyzers
Supported languages:
├─ Hack (primary web language)
├─ C++ (performance-critical services)
├─ Python (ML pipelines)
├─ Java (legacy services)
└─ JavaScript (frontend)
Analysis scope:
├─ Code analyzed: 100M+ lines
├─ Analysis time: Minutes (incremental)
├─ AST nodes: Billions
└─ Data flows found: 100M+
3. SQL Analyzer
Supported engines:
├─ Presto
├─ Spark
├─ Hive
├─ Custom SQL variants
Capabilities:
├─ Table lineage ✅
├─ Column lineage ✅
├─ Transformation tracking ✅
├─ CTE support ✅
├─ Subquery analysis ✅
└─ Window functions ✅
4. Lineage Graph Database
Requirements:
├─ Store: 10M+ nodes, 100M+ edges
├─ Query: Sub-second for most queries
├─ Update: Real-time (1,000+ updates/sec)
├─ Traverse: Multi-hop paths efficiently
└─ Scale: Horizontally
Technology:
├─ Custom graph database
├─ Distributed across data centers
├─ Replicated for reliability
└─ Optimized for graph traversal
Results and Impact
By The Numbers
Coverage:
✅ Assets tracked: 10,000,000+
✅ Data flows mapped: 100,000,000+
✅ Languages supported: 10+
✅ Systems instrumented: 100+
✅ Teams using: 500+
Accuracy:
✅ Precision: 90%+ (match-set)
✅ Recall: 85%+ (combined static + runtime)
✅ False positive rate: <10%
✅ False negative rate: <15%
Performance:
✅ Lineage query time: <30 seconds
✅ Real-time updates: <5 minute delay
✅ System overhead: <1% latency
✅ Cost per query: $0 (after infrastructure)
Time Savings
Privacy control implementation:
Before lineage:
- Manual code review: 6-12 months
- Engineering cost: $5M per requirement
- Error rate: 40-50%
- Coverage: 60-70%
After lineage:
- Automated discovery: 1 day
- Engineering cost: $50K per requirement
- Error rate: <10%
- Coverage: 90%+
Time saved: 99%
Cost saved: 99%
Quality improved: 4x ✅
Privacy Impact
Privacy controls deployed:
├─ Purpose limitations: 1,000+
├─ Data flows protected: 100,000+
├─ Privacy violations prevented: 100s per year
└─ Compliance verified: Continuous
Regulatory compliance:
✅ GDPR: Auditable compliance
✅ CCPA: Verified enforcement
✅ COPPA: Automatic children's data protection
✅ Others: Framework extensible
Developer Productivity
Before:
├─ 70% time on data flow discovery
├─ 20% time on privacy reviews
└─ 10% time on feature development
After:
├─ 10% time on data flow discovery (automated)
├─ 20% time on privacy reviews (streamlined)
└─ 70% time on feature development ✅
Developer happiness: Significantly improved
Product velocity: 7x faster for privacy work
Key Learnings and Challenges
Learning 1: Focus on Lineage Early
Insight: Data lineage should be the FIRST step in privacy infrastructure, not an afterthought.
What Meta learned:
Initial approach:
1. Build Policy Zones (privacy controls)
2. Manually find where to apply them
3. Realize: Can't scale without lineage
4. Build lineage
5. Accelerate Policy Zones adoption
Better approach:
1. Build data lineage first
2. Use lineage to guide Policy Zones rollout
3. Much faster adoption
4. Better coverage
Lesson: Invest in lineage before building controls
Learning 2: Build Consumption Tools
Insight: Raw lineage data is useless without great tools.
What went wrong:
Year 1:
- Built lineage collection
- Engineers: "Great, where's the query tool?"
- Team: "Just query the database directly"
- Engineers: "This is too complex..."
- Adoption: 10%
Year 2:
- Built Policy Zone Manager (PZM)
- Visual interface
- Interactive exploration
- Bulk operations
- Adoption: 80% ✅
Lesson: Invest 50% in collection, 50% in tools
Learning 3: Integrate with Systems
Insight: Don’t ask every team to instrument their code — provide libraries that do it automatically.
What worked:
Bad approach:
- "Every team: Please add lineage tracking to your code"
- Result: 30% compliance, inconsistent quality
Good approach:
- Instrument core frameworks (logging, DB, APIs)
- All code using frameworks gets lineage automatically
- Result: 90% coverage, consistent quality ✅
Lesson: Make lineage automatic, not optional
Learning 4: Combine Static and Runtime
Insight: Neither static nor runtime analysis alone is sufficient — you need both.
Why both:
Static alone:
✅ Finds all possible flows
❌ Many false positives
❌ Can't see transformations
❌ Can't verify actual behavior
Runtime alone:
✅ High accuracy
❌ Only sees executed code
❌ Sampling gaps
❌ Expensive at scale
Combined:
✅ High recall (static finds everything)
✅ High precision (runtime verifies)
✅ Best accuracy
✅ Practical at scale
Learning 5: Measure Coverage
Insight: You can’t improve what you don’t measure.
What Meta measures:
Coverage metrics:
├─ % of assets with lineage
├─ % of code paths analyzed
├─ % of systems instrumented
├─ % of data flows captured
└─ Tracked weekly, improved continuously
Quality metrics:
├─ Precision (false positive rate)
├─ Recall (false negative rate)
├─ Latency (query response time)
└─ Adoption (teams using the tool)
Result: Continuous improvement
Challenges and Limitations
Challenge 1: Dynamic and Transformed Data
The problem:
# Original data
location = "Home, 123 Main St"
# Transformation 1: Encoding
location_encoded = geocode(location) # lat: 40.7, lon: -74.0
# Transformation 2: Aggregation
location_count = count_unique([location_encoded, ...]) # 42
# Transformation 3: ML embedding
location_vector = model.embed(location) # [0.23, -0.15, 0.67, ...]
Question: Is location_vector still "location data"?
Current approach:
High confidence (automatic):
✅ Direct copies
✅ Substring matching
✅ Simple renames
Low confidence (human review):
⚠️ Encoded values
⚠️ Aggregations
⚠️ ML embeddings
⚠️ Complex transformations
Limitation: Some transformations require human judgment
Challenge 2: Scale and Performance
The trade-offs:
Sampling rate vs accuracy:
├─ 100% sampling: Perfect accuracy, 10x cost
├─ 10% sampling: Good accuracy, 2x cost
├─ 1% sampling: Acceptable accuracy, 1x cost
└─ Current: 1% sampling (tuned per system)
Graph size vs query performance:
├─ All assets: Complete, slow queries
├─ Filtered assets: Incomplete, fast queries
└─ Current: Intelligent indexing + caching
Trade-off: Accept some gaps for practical performance
Challenge 3: Constant Evolution
The moving target:
Changes per day:
├─ New code: 1,000+ commits
├─ New tables: 100s
├─ New services: 10s
├─ Schema changes: 100s
└─ Always updating
Challenge:
- Lineage must update in real-time
- Can't afford batch recomputation
- Must handle incremental updates
- Need to detect and alert on new flows
Solution:
- Incremental updates
- Event-driven processing
- Continuous monitoring
- Alert on policy violations
Challenge 4: False Positives
The review burden:
At Meta scale:
├─ Total flows: 100M+
├─ False positive rate: 10%
├─ False positives: 10M flows
└─ Human review: Impossible for all
Strategy:
1. Auto-exclude obvious false positives (heuristics)
2. Focus review on high-confidence flows
3. Iterative filtering (exclude cascades)
4. ML models to predict false positives
5. Continuous improvement
Result: Manageable review load (100s, not millions)
Future Directions
Expansion: More Coverage
Expanding lineage to:
├─ Mobile apps (iOS, Android)
├─ Edge computing
├─ Third-party integrations
├─ Encrypted data
└─ Cross-platform flows
Goal: 99%+ coverage of all data flows
Improvement: Better Accuracy
Improving transformation tracking:
├─ ML models to match encoded data
├─ Semantic similarity for embeddings
├─ Pattern recognition for aggregations
└─ Automated confidence scoring
Goal: 95%+ precision and recall
Innovation: New Use Cases
Beyond privacy:
├─ Security: Track sensitive data for security
├─ Integrity: Prevent data quality issues
├─ Compliance: Audit trails for regulations
├─ Cost optimization: Identify unused data
└─ Data governance: Catalog and discovery
Potential: Lineage as universal platform
How You Can Apply This
For Small Companies (10-100 engineers)
Start simple:
├─ Document critical data flows manually
├─ Add lightweight instrumentation
├─ Use open-source lineage tools (e.g., OpenLineage)
├─ Focus on most sensitive data first
└─ Build incrementally
Investment: $100K-$500K
Timeline: 6-12 months
Result: Good enough for compliance
For Mid-Size Companies (100-1,000 engineers)
Build basic automation:
├─ Static analysis for code lineage
├─ SQL parsing for data warehouse
├─ Runtime logging for critical paths
├─ Simple graph database
└─ Internal query tool
Investment: $1M-$5M
Timeline: 1-2 years
Result: Automated lineage for core systems
For Large Companies (1,000+ engineers)
Full automation like Meta:
├─ Multi-language static analysis
├─ Runtime instrumentation framework
├─ Distributed graph database
├─ Developer-friendly tools
├─ Continuous monitoring
└─ ML for transformation tracking
Investment: $10M-$50M
Timeline: 2-4 years
Result: Complete automated lineage at scale
Open Source Options
Available tools:
├─ OpenLineage (LFAI & Data)
│ └─ Standard for lineage metadata
│
├─ Apache Atlas
│ └─ Data governance and lineage
│
├─ Marquez (WeWork)
│ └─ Metadata service for lineage
│
├─ Amundsen (Lyft)
│ └─ Data discovery with lineage
│
└─ DataHub (LinkedIn)
└─ Metadata platform with lineage
Start here before building custom!
Conclusion
The Big Picture
Meta's data lineage journey:
The problem:
├─ Can't protect privacy without knowing where data goes
├─ Manual tracking breaks down at scale
├─ Billions of users, millions of assets
└─ Regulatory requirements (GDPR, CCPA, etc.)
The solution:
├─ Automated lineage collection (static + runtime)
├─ Privacy Probes for accurate flow tracking
├─ SQL analysis for batch systems
├─ Developer-friendly tools (PZM)
└─ Continuous monitoring
The results:
├─ 10M+ assets tracked
├─ 100M+ data flows mapped
├─ 99% time saved on privacy work
├─ 90%+ accuracy
├─ Enabled Privacy Aware Infrastructure at scale
└─ Compliance with global regulations
The impact:
✅ Billions of users protected
✅ Privacy violations prevented
✅ Developer productivity 7x improved
✅ Regulatory compliance verified
✅ Product innovation enabled
Key Takeaways
- Privacy at scale requires automation
- Manual tracking impossible beyond 100 engineers
- Need infrastructure, not policies
- Lineage is the foundation
- Combine static and runtime analysis
- Static: High recall, finds all possible flows
- Runtime: High precision, verifies actual flows
- Together: Best accuracy
- Invest in developer tools
- Raw lineage data is useless
- Need visual, interactive tools
- PZM reduced discovery time by 99%
- Start with core frameworks
- Don’t ask teams to instrument code
- Instrument core libraries
- Automatic coverage
- Measure and improve continuously
- Track coverage metrics
- Monitor accuracy
- Iterate on quality
The Future
Data lineage is becoming essential infrastructure — not just for privacy, but for security, compliance, cost optimization, and data governance.
As data ecosystems grow more complex, lineage will evolve from a “nice to have” to a fundamental requirement for operating at scale.
Further Reading
- Original Meta Engineering Blog - Source article with more technical details
- Policy Zones at Meta - How Meta enforces purpose limitation
- Privacy Aware Infrastructure - PAI overview
- OpenLineage Project - Open standard for lineage metadata
- Data Governance at Scale - Broader context on data governance
- GDPR Purpose Limitation - Regulatory requirements