Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #232Beginner 8 min read

Design a Pastebin Service

๐Ÿ’ก
Core Architecture Summary

Store and share plain text snippets: Object storage (S3) for paste content, Relational/NoSQL metadata indexing, TTL auto-expiration, and custom vanity URLs.

Key Glossary Concepts in this TopicAll Glossary Terms

Pastebin Decoupled Metadata & Object Storage Architecture ๐Ÿ“‹

Separation of lightweight metadata in PostgreSQL/DynamoDB from bulk text payloads in Amazon S3, fronted by Redis and Cloudflare CDN.

Pastebin Decoupled Metadata & Object Storage Architecture ๐Ÿ“‹
100%
Rendering visual architecture flowchart...

01.1. Functional & Non-Functional Requirements

A Pastebin service allows users to store plain text or source code snippets and generate shareable short URLs with optional expiration timestamps.

Functional Requirements

  1. Paste Creation: Users can paste text up to 10 MB per snippet and receive a unique 7-character URL (e.g., https://pastebin.com/k9Xa1Zb).
  2. Paste Retrieval: Anyone with the short link can view the raw text or syntax-highlighted snippet.
  3. Expiration (TTL): Pastes can be configured to expire after 1 hour, 1 day, 1 week, 1 month, or Never.
  4. Access Control: Public pastes, unlisted pastes (accessible only via link), and private password-protected pastes.
  5. Custom Vanity Slugs: Optional custom URL paths (e.g., https://pastebin.com/my-error-log).

Non-Functional Requirements

  • High Availability & Durability: 99.99\% uptime. Once written, pastes must be durably stored across multiple availability zones.
  • Low Read Latency: Retrieve cached pastes in < 10 ms (p99) and un-cached pastes in < 50 ms.
  • Cost Efficiency at Scale: Store petabytes of text cost-effectively without bloating relational database memory buffers.

02.2. Capacity & Scale Calculations

Traffic Estimates

  • Write Traffic: 10 million new pastes per day.

Write QPS = \frac{10,000,000}{86,400} โ‰ˆ 116 writes/sec (Peak: 300 writes/sec)

  • Read Traffic: 10:1 Read-to-Write ratio \implies 100 million reads per day.

Read QPS = \frac{100,000,000}{86,400} โ‰ˆ 1,160 reads/sec (Peak: 4,000 reads/sec)

Storage Calculations (3-Year Retention)

  • Average Paste Size: 10 KB (Text files average 10KB; max limit 10MB).
  • Daily Storage Ingestion:

10M pastes/day ร— 10 KB = 100 GB/day

  • 3-Year Persistent Storage:

100 GB/day ร— 365 ร— 3 โ‰ˆ 109.5 Terabytes (TB)

In-Memory Caching (80/20 Rule)

  • 20% of the daily read traffic (20M reads) targets popular pastes:

20M pastes ร— 10 KB = 200 GB of RAM

A small Redis cluster of 4 instances with 64GB RAM each comfortably holds the entire hot working set.

03.3. Storage Architecture: Decoupling Metadata and Raw Text Blobs

A classic anti-pattern in Pastebin design is storing 10MB text strings directly inside PostgreSQL TEXT or BLOB columns. Large blobs cause massive disk I/O, fragment database pages, and exhaust database buffer pool RAM.

The Decoupled Two-Tier Architecture:

  1. Metadata Store (Relational PostgreSQL or DynamoDB): Stores lightweight index records (< 1 KB per row).
  2. Object Store (AWS S3 / Google Cloud Storage / Ceph): Stores raw text files as immutable objects named by their paste ID (s3://pastes-bucket/k9Xa1Zb.txt).
sql
-- Metadata Table DDL (PostgreSQL)
CREATE TABLE pastes (
    paste_id VARCHAR(10) PRIMARY KEY,
    user_id UUID,
    title VARCHAR(255),
    syntax_language VARCHAR(50),
    s3_object_key VARCHAR(255) NOT NULL,
    size_bytes BIGINT NOT NULL,
    access_type VARCHAR(20) DEFAULT 'PUBLIC', -- PUBLIC, UNLISTED, PRIVATE
    password_hash VARCHAR(255),
    created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
    expires_at TIMESTAMP WITH TIME ZONE
);

CREATE INDEX idx_pastes_user_id ON pastes(user_id);
CREATE INDEX idx_pastes_expires_at ON pastes(expires_at) WHERE expires_at IS NOT NULL;

04.4. API Design & Endpoints

Endpoints

  1. Create a Paste

    • POST /api/v1/pastes
    • Request Body:
      json
      {
        "content": "SELECT * FROM users WHERE active = true;",
        "title": "SQL Query Snippet",
        "syntax": "sql",
        "ttl_seconds": 86400,
        "access": "PUBLIC"
      }
    • Response (201 Created):
      json
      {
        "paste_id": "k9Xa1Zb",
        "url": "https://pastebin.com/k9Xa1Zb",
        "expires_at": "2026-09-28T10:00:00Z"
      }
  2. Retrieve Paste Metadata & Content

    • GET /api/v1/pastes/{paste_id}
    • Response (200 OK):
      json
      {
        "paste_id": "k9Xa1Zb",
        "title": "SQL Query Snippet",
        "syntax": "sql",
        "content": "SELECT * FROM users WHERE active = true;",
        "created_at": "2026-09-27T10:00:00Z"
      }

05.5. Handling Expired Pastes & Automated Purging

Expired pastes must be reliably purged from both the database and S3 to reclaim storage and prevent zombie data leaks:

  1. Passive Deletion (Lazy Eviction):

    • When a user requests GET /k9Xa1Zb, the service checks expires_at < NOW().
    • If expired, return HTTP 404 Not Found immediately and enqueue an asynchronous deletion task.
  2. Active Scheduled Sweeper (Background Worker):

    • A cron job runs every hour querying SELECT paste_id, s3_object_key FROM pastes WHERE expires_at < NOW() LIMIT 5000;.
    • The worker batch-deletes objects from S3 via the S3 DeleteObjects API and deletes the database metadata rows.
  3. S3 Object Lifecycle Policies:

    • Configure S3 bucket lifecycle rules using object tags (expire_date) to automate cloud-level blob deletion without running custom worker clusters.

06.6. Performance Optimizations & Security

Edge CDN Caching

  • Public pastes that are frequently accessed (e.g., viral error logs or open-source configurations) are cached at Cloudflare/CloudFront edge nodes for the duration of their TTL.

Abuse Prevention & Malicious Content Scanning

  • Pastebin services are primary targets for malicious payload hosting (malware droppers, phishing scripts, stolen API keys).
  • Asynchronous Security Pipeline: Every paste write enqueues a message to Kafka. A background security worker scans text using ClamAV and regex DLP filters to identify and quarantine leaked AWS credentials, credit card numbers, and malware strings.

โš–๏ธArchitectural Trade-offs & Production Realities

Architectural Advantages

  • Separating metadata from S3 object storage keeps database buffer pools lean and reduces storage costs by 90%
  • Redis and Edge CDN absorb over 95% of read traffic, delivering sub-10ms response times
  • S3 lifecycle policies and lazy deletion automate multi-terabyte data expiration

Trade-offs & Constraints

  • Cache miss on read requires a two-step lookup (metadata DB query followed by S3 blob fetch)
  • Large text pastes (10MB) can saturate server network cards without direct presigned S3 uploads
Production Implementation in Big Tech
GitHub Gistsโ€ข Code Snippet Storage & Rendering

GitHub Gists stores snippet metadata in sharded MySQL clusters and git file blobs in distributed block/object storage, serving cached syntax-highlighted renders through Fastly CDN edge nodes.

๐ŸŽฏ Staff+ Engineering Takeaways

  • Store lightweight metadata in a database and bulky text payloads in S3 object storage.
  • Use Base62 7-character IDs generated via Range-Based Counters for short URLs.
  • Cache hot pastes in Redis to achieve sub-10ms read latencies.
  • Implement dual expiration: lazy deletion on read and scheduled background S3 sweeping.

Topic Knowledge Assessment ๐Ÿง 

Step through 2 scenario questions to test your staff-level grasp.

Question 1 of 20 answered
#1

Why should the raw 5MB text body of a paste be stored in AWS S3 rather than directly inside a relational PostgreSQL table column?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?