Meta Data Lineage

Part 1 of the Meta Data Lineage Series

Important Note: This article is based on my understanding after reading the Meta Engineering blog post and several related articles. I’m trying to demystify and explain these concepts in an accessible way. If you want to understand exactly what Meta built, please refer to the original article linked in the Further Reading section.

The surprising problem: Meta has billions of users and world-class infrastructure. But when someone enters their religion on Facebook Dating, can they guarantee it won’t be used elsewhere? The answer required building something unprecedented.

This article explains why data lineage became critical for privacy at Meta’s scale — and why manual tracking completely breaks down.

Next: Part 2 - The Solution (Building Data Lineage at Scale) →

Introduction: A Simple Privacy Promise

Imagine you’re sharing a family photo on Facebook - a picture of your kids at home:

User shares photo:
📷 Family photo with location: "Home address: 123 Main St"

User's expectation:
✅ Visible to friends on Facebook
✅ Maybe suggest nearby friends
❌ NOT used to target ads based on home location
❌ NOT shared with third-party apps
❌ NOT used to track family members' movements
❌ NOT visible on other Meta apps without permission

Question: Can Meta guarantee this?

For a small app: This is easy. You know where all your data goes.

For Meta (billions of users, millions of code assets): This is one of the hardest technical challenges imaginable.

Think of it like this: Imagine a city with millions of water pipes, and you need to guarantee that water from one specific faucet never reaches certain buildings. Now imagine the pipes are constantly being rebuilt while water is flowing.

This is the story of how Meta built data lineage — the technology that traces data as it flows through their massive systems — to enable Privacy Aware Infrastructure (PAI) at unprecedented scale.


What is Privacy Aware Infrastructure (PAI)?

The Simple Explanation

Privacy Aware Infrastructure (PAI) is Meta’s suite of technologies that enforce privacy rules automatically in code, rather than relying on humans to follow guidelines.

The core principle:

Old approach (doesn't scale):
Write policies → Train developers → Hope they follow rules → Audit manually

PAI approach (scales):
Write policies → Embed in infrastructure → Impossible to violate → Automatic verification

Why PAI is Needed

The regulatory landscape:

2018: GDPR in Europe
- Requires "purpose limitation"
- Data can only be used for stated purposes
- Violations: Up to 4% of global revenue ($4.6B+ for Meta)

2020: CCPA in California
- Similar requirements
- Growing globally

2023: 100+ privacy laws worldwide
- Each with different requirements
- Manual compliance: impossible at scale

The scale problem:

Meta's challenge:
├─ 3.8 billion users
├─ 100+ products (Facebook, Instagram, WhatsApp, etc.)
├─ Millions of data assets (tables, APIs, models)
├─ Billions of lines of code
├─ Thousands of engineers
└─ Code changes: 1,000+ per day

Manual privacy tracking:
❌ Impossible
❌ Error-prone
❌ Doesn't scale
❌ Can't verify

Example: Purpose Limitation

Purpose limitation restricts the purposes for which data can be processed and used.

Real example:

User shares photo with location: "At Disneyland with kids"

Allowed uses (explicit consent):
✅ Show photo to friends
✅ Suggest friends nearby
✅ Add to family memories

NOT allowed (no consent):
❌ Target location-based ads to the user
❌ Track children's locations
❌ Share with third-party advertisers
❌ Use for surveillance purposes

Challenge: How do you enforce this across billions of lines of code?

The Data Lineage Problem

What is Data Lineage?

Data lineage traces the journey of data as it moves through systems, from source to sink.

Simple example:

User shares photo with location (SOURCE)
    ↓
Web Endpoint
    ↓
Database Table
    ↓
Data Warehouse
    ↓
AI Model
    ↓
Ad Targeting System (SINK)

Question: Does location data from photos reach the ad targeting system?

Why Data Lineage is Critical for Privacy

The PAI workflow:

flowchart TD START["1. Inventory Assets
(catalog all data)"] SCHEMA["2. Schematize
(define structure)"] ANNOTATE["3. Annotate
(label sensitive data)"] LINEAGE["4. Data Lineage
(trace data flows)"] CONTROL["5. Privacy Controls
(enforce policies)"] VERIFY["6. Verification
(continuous monitoring)"] START --> SCHEMA SCHEMA --> ANNOTATE ANNOTATE --> LINEAGE LINEAGE --> CONTROL CONTROL --> VERIFY VERIFY -.-> LINEAGE style LINEAGE fill:#fef3c7,stroke:#f59e0b style CONTROL fill:#dbeafe,stroke:#2563eb

Data lineage enables:

  1. Discovery: “Where does this sensitive data go?”
  2. Control placement: “Where should we add privacy controls?”
  3. Verification: “Are our controls actually working?”

Without lineage, you’re flying blind.


The Scale of the Challenge

Meta’s Data Landscape

Assets to track:
├─ Web endpoints: 100,000+
├─ Database tables: 1,000,000+
├─ Data warehouse tables: 10,000,000+
├─ AI models: 100,000+
├─ Logging systems: 10,000+
└─ Microservices: 50,000+

Technology stacks:
├─ Languages: Hack, C++, Python, JavaScript, Rust, Java
├─ Data systems: MySQL, RocksDB, Presto, Spark, Hive
├─ Processing: Web servers, batch jobs, streaming pipelines
└─ ML platforms: PyTorch, TensorFlow, custom frameworks

Code changes:
├─ 1,000+ commits per day
├─ 5,000+ engineers making changes
├─ Assets constantly being created/modified/deleted
└─ No way to manually track

The Problem: Manual Tracking Doesn’t Scale

Traditional approach (what most companies do):

Manual data lineage:
├─ Engineers document data flows
├─ Create diagrams in slides/wikis
├─ Update spreadsheets
└─ Hope it stays accurate

Reality:
❌ Documentation immediately out of date
❌ Engineers forget to update
❌ Takes weeks to trace one data flow
❌ Incomplete (only ~10% coverage)
❌ Can't verify accuracy

Example: Tracing location data manually

Week 1: Start investigation
- Interview Photos team: "Where does location data from photos go?"
- Photos team: "Uh... to the database, some logs, maybe some APIs?"
- Manually read code to find all references

Week 2: Trace database
- Find 67 downstream tables that read from photo_metadata
- Each has 20-100 downstream consumers
- Need to check each one

Week 3: Trace APIs
- Find 25 APIs that expose photo/location data
- Who calls these APIs? Need to search all code
- Each API has 50+ callers

Week 4: Trace logs
- Location data might be in 300+ log tables
- Each log feeds into data warehouse
- Data warehouse has 1,000s of downstream tables

Week 8: Still not done
- Found 10,000+ potential data flows
- 90% are false positives
- Still missing actual flows
- Code has changed since week 1 (already out of date!)

Result: 
- 2 months spent
- Incomplete and inaccurate
- Out of date immediately
- Costs $50,000+ per investigation
- Impossible to scale

Why Manual Tracking Fails

Problem 1: Dynamic systems

While you're documenting:
├─ New tables are created
├─ New code is deployed
├─ New data flows are added
└─ Documentation is already wrong

It's like mapping a city where streets
change every minute!

Problem 2: Transformed data

Input: location = "Disneyland, Anaheim, CA"

Transformation 1:
location → {metadata: {location: "Disneyland, Anaheim, CA", lat: 33.8, lon: -117.9}}

Transformation 2:
{locations: ["Disneyland", "Universal Studios"]} → {nearby_count: 2}

Transformation 3:
"Disneyland, Anaheim, CA" → "LA_AREA_THEME_PARK" (category code)

Question: Are these the same data?
Manual tracking: Impossible to tell

Problem 3: Indirect flows

Direct flow (easy to track):
photo.location → nearby_friends_feature

Indirect flow (nearly impossible to track):
photo.location 
  → tmp_table_abc123
  → aggregated_location_features
  → model_training_dataset
  → user_interest_model
  → ad_targeting_service
  → third_party_advertisers

Manual tracking: Would take months to discover this chain

Problem 4: Scale

To manually track location data:
├─ Review: 1,000,000+ lines of code
├─ Interview: 500+ engineers
├─ Check: 100,000+ assets
├─ Trace: 1,000,000+ potential flows
└─ Time: 6-12 months (minimum)

But:
├─ Code changes daily
├─ Engineers leave/join
├─ Systems evolve
└─ Impossible to keep up

Cost: $5M+ per privacy requirement
Timeline: Months to years
Accuracy: 50-60% at best

Real Consequences of Not Having Lineage

Consequence 1: Can’t Implement Privacy Controls

The challenge:

GDPR requirement: Location data must be purpose-limited

To implement:
1. Find all places location data is used
2. Add privacy controls at those places
3. Verify controls work

Without lineage:
❌ Can't find all places (manual search takes months)
❌ Miss hidden flows (transformed data)
❌ Can't verify (no visibility)

Result: Can't comply with GDPR
Risk: $4.6B fine (4% of global revenue)

Consequence 2: Privacy Violations

Real scenario:

2018: Developer adds new feature

Code:
def show_targeted_ads(user_id):
    photos = get_user_photos(user_id)
    # Use ALL photo metadata including location
    return target_ads_by_location(photos)

Oops: Just used children's location data for ad targeting!

Without lineage:
- Nobody knows location data leaked to ads
- No alerts
- No way to discover
- Privacy violation goes undetected

With lineage:
- Automatic detection: "Photo location data flowing to ad targeting"
- Alert: "Potential privacy violation - children's data"
- Block: "Cannot deploy code that violates policy"

Consequence 3: Slow Product Development

The bottleneck:

Product team: "We want to add a new feature"

Privacy team: "First, we need to verify data flows"

Manual review:
- Week 1-4: Document current data flows
- Week 5-8: Trace new data flows
- Week 9-12: Write privacy review
- Week 13-16: Get approvals
- Week 17-20: Implement controls

Total: 5 months before feature can launch

Result:
- Product team frustrated
- Innovation slowed
- Competitors move faster
- Company loses market share

Consequence 4: Can’t Audit or Verify

The regulatory problem:

Auditor: "How do you ensure religion data from Dating 
          doesn't reach your ad system?"

Without lineage:
Company: "We have policies and training"
Auditor: "Can you prove it?"
Company: "Uh... we trust our engineers?"
Auditor: "That's not acceptable"
Company: "We can review code manually?"
Auditor: "How long would that take?"
Company: "6-12 months... per data type..."
Auditor: "Not acceptable. You're not compliant."

Result: Can't pass audit, can't operate

Why Traditional Solutions Don’t Work

Failed Solution 1: “Just Document Everything”

The approach:

Create massive documentation:
├─ Data dictionaries
├─ Flow diagrams
├─ Spreadsheets
└─ Wikis

Reality:
- Documentation created: 10,000 pages
- Time to create: 1 year
- Cost: $2M
- Out of date: Immediately
- Engineers read it: 5%
- Actual use: Near zero

Result: $2M wasted ❌

Failed Solution 2: “Manual Code Reviews”

The approach:

Require privacy review for every code change

Reality:
- Code changes per day: 1,000+
- Time per review: 2 hours
- Reviewers needed: 250 full-time
- Cost: $50M per year
- Bottleneck: Everything slows down
- Engineers bypass process

Result: Creates bureaucracy, doesn't scale ❌

Failed Solution 3: “Privacy Training”

The approach:

Train all engineers on privacy rules

Reality:
- Engineers trained: 5,000
- Time per training: 4 hours
- Cost: $2M per year
- Engineers remember: 20%
- Engineers follow rules: 30%
- Mistakes still happen: Daily

Result: Doesn't prevent violations ❌

Failed Solution 4: “Ask Engineers Where Data Goes”

The approach:

"Just ask the team that owns the code"

Reality:
Engineer: "Uh, I think it goes to these 5 tables..."
[Actually goes to 50+ tables]

Engineer: "I don't know, I didn't write this code"
[Original author left 2 years ago]

Engineer: "It's in the logs somewhere"
[Logs go to 100+ downstream systems]

Result: Incomplete, inaccurate ❌

What Meta Realized

The Key Insight

You can’t enforce privacy without understanding data flows.

And you can’t understand data flows at Meta’s scale without automation.

The realization:

Old thinking:
"Privacy is about policies and training"

New thinking:
"Privacy is about infrastructure and automation"

Old approach:
- Write privacy policies
- Train developers
- Review code manually
- Hope nothing breaks

New approach (PAI):
- Build privacy into infrastructure
- Automatic data lineage
- Automatic control enforcement
- Automatic verification

The Requirements

What Meta needed to build:

1. Automated data lineage collection
   ✅ Zero manual work
   ✅ Tracks all data flows
   ✅ Updates in real-time
   ✅ Handles all tech stacks

2. High accuracy
   ✅ 90%+ precision
   ✅ Minimal false positives
   ✅ Catches transformations
   ✅ Finds indirect flows

3. Complete coverage
   ✅ All programming languages
   ✅ All data systems
   ✅ All processing types
   ✅ Millions of assets

4. Developer-friendly
   ✅ Easy to query
   ✅ Visual tools
   ✅ Fast results
   ✅ Actionable insights

5. Scalable
   ✅ Handles billions of data flows
   ✅ Processes in real-time
   ✅ Cost-effective
   ✅ Reliable

6. Verifiable
   ✅ Can prove compliance
   ✅ Audit trail
   ✅ Continuous monitoring
   ✅ Alerts for violations

The Challenge

Building this is incredibly hard:

Technical challenges:
├─ Multiple programming languages
├─ Different data paradigms (SQL, APIs, models)
├─ Transformed data (hard to match)
├─ Millions of assets
├─ Billions of data flows
├─ Real-time updates
├─ Minimal performance impact
└─ Must be 90%+ accurate

Organizational challenges:
├─ 5,000+ engineers to support
├─ Can't disrupt existing systems
├─ Must integrate with all tech stacks
├─ Gradual rollout (can't migrate everything at once)
└─ Adoption across hundreds of teams

Timeline: 3+ years
Investment: $50M+
Team size: 100+ engineers

Examples: Where Data Lineage is Critical

Example 1: Family Photo Location Data

The requirement:

User shares family photo with location: "Home, 123 Main St"

Allowed uses:
✅ Show photo to friends
✅ Create family memories album
✅ Suggest nearby friends once

Not allowed:
❌ Target ads based on home location
❌ Track family/children's locations over time
❌ Share with third-party apps
❌ Use for location-based marketing

Question: How do you enforce this across millions of lines of code?
Answer: You need data lineage!

Without lineage:

Engineer adds new ad targeting feature:
def target_local_ads(user_id):
    photos = get_user_photos(user_id)  # Oops! Includes home location
    locations = extract_locations(photos)  # Now location leaked
    return show_nearby_business_ads(locations)

Result: Privacy violation, nobody knows ❌

With lineage:

Data lineage system detects:
- Photo location data flowing from Photos → ad targeting
- Automatic alert: "PRIVACY VIOLATION DETECTED"
- Block deployment: "Cannot deploy code that violates policy"
- Alert teams: "Fix privacy violation before deploying"

Result: Privacy protected automatically ✅

Example 2: Children’s Location Data

The requirement:

COPPA (Children's Online Privacy Protection Act):
- Data from users under 13 has strict restrictions
- Location data especially sensitive
- Cannot be used for behavioral advertising
- Cannot be shared with third parties
- Cannot be used to track children

Example: Kid posts photo from school
- School location must be protected
- Cannot be used for ads
- Cannot be shared
- Must be deleted after X days

Challenge: Need to track all flows of children's location data
across all Meta products

Scale:
- 100M+ children on Meta platforms
- Billions of photos with locations
- Data in millions of assets
- Need 100% accuracy (regulatory requirement)

Manual tracking: Impossible
Automated lineage: Essential

The Bottom Line

Why Data Lineage Matters

For privacy:

  1. Enables compliance
    • GDPR, CCPA, COPPA, etc.
    • Can’t comply without knowing where data goes
    • Violations: $4.6B+ in potential fines
  2. Protects users
    • Ensures data used only as promised
    • Prevents unauthorized use
    • Builds trust
  3. Enables verification
    • Can prove controls work
    • Pass audits
    • Continuous monitoring
  4. Speeds product development
    • Automatic privacy checks (seconds vs months)
    • No bottlenecks
    • Innovation enabled

The Stakes

Without data lineage:
├─ Can't implement privacy controls
├─ Can't verify compliance
├─ Risk: $4.6B+ in fines
├─ User trust destroyed
└─ Product development slowed

With data lineage:
├─ Automatic privacy enforcement
├─ Continuous compliance verification
├─ Risk: Minimized
├─ User trust maintained
└─ Fast product iteration

Data lineage is not optional—it's essential

What’s Next?

In Part 2, we’ll see how Meta built their data lineage system:

Sneak peek of the solution:

Meta's data lineage system:

Collection methods:
├─ Static code analysis (simulates execution)
├─ Privacy Probes (runtime instrumentation)
├─ SQL analysis (batch processing)
├─ Config parsing (AI systems)
└─ Integrated across all tech stacks

Key technologies:
├─ Automatic payload matching
├─ Cross-language support
├─ Real-time processing
├─ Graph database for relationships
└─ Developer-friendly query tools

Results:
✅ Millions of assets tracked
✅ Billions of data flows mapped
✅ 90%+ accuracy
✅ Real-time updates
✅ Developer-friendly
✅ Enabled PAI at scale

Engineering time saved: 90%+
Privacy violations prevented: 100s per year
Cost: $50M investment, $500M+ value

Continue to: Part 2 - Building Data Lineage at Scale →


Summary: The Data Lineage Challenge

The Problem in One Picture

The Question: "Where does location data from photos go?"

Without data lineage:
├─ Manual code review: 6 months
├─ Incomplete results: 60% coverage
├─ Inaccurate: 50+ false positives per true flow
├─ Out of date: Immediately
└─ Cost: $5M per investigation

With data lineage:
├─ Automated query: 30 seconds
├─ Complete results: 95%+ coverage
├─ Accurate: 90%+ precision
├─ Always up to date: Real-time
└─ Cost: $0 per query (after building system)

Key Takeaways

  1. Privacy at scale requires data lineage
    • Can’t protect data without knowing where it goes
    • Manual tracking completely breaks down
    • Automation is essential
  2. The stakes are massive
    • Regulatory fines: $4.6B+
    • User trust: Priceless
    • Product velocity: Months vs seconds
  3. Traditional solutions don’t work
    • Documentation, training, reviews all fail
    • Need infrastructure, not policies
  4. The scale is unprecedented
    • Millions of assets
    • Billions of data flows
    • Thousands of engineers
    • Changes every minute
  5. Meta had to build it
    • No existing solution
    • 3+ years of development
    • 100+ engineers
    • $50M+ investment

Next: Learn exactly how they built it! → Part 2


Further Reading