Design a Pastebin Service
Store and share plain text snippets: Object storage (S3) for paste content, Relational/NoSQL metadata indexing, TTL auto-expiration, and custom vanity URLs.
Pastebin Decoupled Metadata & Object Storage Architecture ๐
Separation of lightweight metadata in PostgreSQL/DynamoDB from bulk text payloads in Amazon S3, fronted by Redis and Cloudflare CDN.
01.1. Functional & Non-Functional Requirements
A Pastebin service allows users to store plain text or source code snippets and generate shareable short URLs with optional expiration timestamps.
Functional Requirements
- Paste Creation: Users can paste text up to
10 MBper snippet and receive a unique 7-character URL (e.g.,https://pastebin.com/k9Xa1Zb). - Paste Retrieval: Anyone with the short link can view the raw text or syntax-highlighted snippet.
- Expiration (TTL): Pastes can be configured to expire after
1 hour,1 day,1 week,1 month, orNever. - Access Control: Public pastes, unlisted pastes (accessible only via link), and private password-protected pastes.
- Custom Vanity Slugs: Optional custom URL paths (e.g.,
https://pastebin.com/my-error-log).
Non-Functional Requirements
- High Availability & Durability:
99.99\%uptime. Once written, pastes must be durably stored across multiple availability zones. - Low Read Latency: Retrieve cached pastes in
< 10 ms(p99) and un-cached pastes in< 50 ms. - Cost Efficiency at Scale: Store petabytes of text cost-effectively without bloating relational database memory buffers.
02.2. Capacity & Scale Calculations
Traffic Estimates
- Write Traffic:
10 millionnew pastes per day.
Write QPS = \frac{10,000,000}{86,400} โ 116 writes/sec (Peak: 300 writes/sec)
- Read Traffic: 10:1 Read-to-Write ratio
\implies 100 millionreads per day.
Read QPS = \frac{100,000,000}{86,400} โ 1,160 reads/sec (Peak: 4,000 reads/sec)
Storage Calculations (3-Year Retention)
- Average Paste Size:
10 KB(Text files average 10KB; max limit 10MB). - Daily Storage Ingestion:
10M pastes/day ร 10 KB = 100 GB/day
- 3-Year Persistent Storage:
100 GB/day ร 365 ร 3 โ 109.5 Terabytes (TB)
In-Memory Caching (80/20 Rule)
- 20% of the daily read traffic (
20M reads) targets popular pastes:
20M pastes ร 10 KB = 200 GB of RAM
A small Redis cluster of 4 instances with 64GB RAM each comfortably holds the entire hot working set.
03.3. Storage Architecture: Decoupling Metadata and Raw Text Blobs
A classic anti-pattern in Pastebin design is storing 10MB text strings directly inside PostgreSQL TEXT or BLOB columns. Large blobs cause massive disk I/O, fragment database pages, and exhaust database buffer pool RAM.
The Decoupled Two-Tier Architecture:
- Metadata Store (Relational PostgreSQL or DynamoDB): Stores lightweight index records (
< 1 KBper row). - Object Store (AWS S3 / Google Cloud Storage / Ceph): Stores raw text files as immutable objects named by their paste ID (
s3://pastes-bucket/k9Xa1Zb.txt).
sql-- Metadata Table DDL (PostgreSQL) CREATE TABLE pastes ( paste_id VARCHAR(10) PRIMARY KEY, user_id UUID, title VARCHAR(255), syntax_language VARCHAR(50), s3_object_key VARCHAR(255) NOT NULL, size_bytes BIGINT NOT NULL, access_type VARCHAR(20) DEFAULT 'PUBLIC', -- PUBLIC, UNLISTED, PRIVATE password_hash VARCHAR(255), created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(), expires_at TIMESTAMP WITH TIME ZONE ); CREATE INDEX idx_pastes_user_id ON pastes(user_id); CREATE INDEX idx_pastes_expires_at ON pastes(expires_at) WHERE expires_at IS NOT NULL;
04.4. API Design & Endpoints
Endpoints
-
Create a Paste
POST /api/v1/pastes- Request Body:
json
{ "content": "SELECT * FROM users WHERE active = true;", "title": "SQL Query Snippet", "syntax": "sql", "ttl_seconds": 86400, "access": "PUBLIC" } - Response (201 Created):
json
{ "paste_id": "k9Xa1Zb", "url": "https://pastebin.com/k9Xa1Zb", "expires_at": "2026-09-28T10:00:00Z" }
-
Retrieve Paste Metadata & Content
GET /api/v1/pastes/{paste_id}- Response (200 OK):
json
{ "paste_id": "k9Xa1Zb", "title": "SQL Query Snippet", "syntax": "sql", "content": "SELECT * FROM users WHERE active = true;", "created_at": "2026-09-27T10:00:00Z" }
05.5. Handling Expired Pastes & Automated Purging
Expired pastes must be reliably purged from both the database and S3 to reclaim storage and prevent zombie data leaks:
-
Passive Deletion (Lazy Eviction):
- When a user requests
GET /k9Xa1Zb, the service checksexpires_at < NOW(). - If expired, return HTTP 404 Not Found immediately and enqueue an asynchronous deletion task.
- When a user requests
-
Active Scheduled Sweeper (Background Worker):
- A cron job runs every hour querying
SELECT paste_id, s3_object_key FROM pastes WHERE expires_at < NOW() LIMIT 5000;. - The worker batch-deletes objects from S3 via the S3
DeleteObjectsAPI and deletes the database metadata rows.
- A cron job runs every hour querying
-
S3 Object Lifecycle Policies:
- Configure S3 bucket lifecycle rules using object tags (
expire_date) to automate cloud-level blob deletion without running custom worker clusters.
- Configure S3 bucket lifecycle rules using object tags (
06.6. Performance Optimizations & Security
Edge CDN Caching
- Public pastes that are frequently accessed (e.g., viral error logs or open-source configurations) are cached at Cloudflare/CloudFront edge nodes for the duration of their TTL.
Abuse Prevention & Malicious Content Scanning
- Pastebin services are primary targets for malicious payload hosting (malware droppers, phishing scripts, stolen API keys).
- Asynchronous Security Pipeline: Every paste write enqueues a message to Kafka. A background security worker scans text using ClamAV and regex DLP filters to identify and quarantine leaked AWS credentials, credit card numbers, and malware strings.
โ๏ธArchitectural Trade-offs & Production Realities
Architectural Advantages
- Separating metadata from S3 object storage keeps database buffer pools lean and reduces storage costs by 90%
- Redis and Edge CDN absorb over 95% of read traffic, delivering sub-10ms response times
- S3 lifecycle policies and lazy deletion automate multi-terabyte data expiration
Trade-offs & Constraints
- Cache miss on read requires a two-step lookup (metadata DB query followed by S3 blob fetch)
- Large text pastes (10MB) can saturate server network cards without direct presigned S3 uploads
GitHub Gists stores snippet metadata in sharded MySQL clusters and git file blobs in distributed block/object storage, serving cached syntax-highlighted renders through Fastly CDN edge nodes.
๐ฏ Staff+ Engineering Takeaways
- Store lightweight metadata in a database and bulky text payloads in S3 object storage.
- Use Base62 7-character IDs generated via Range-Based Counters for short URLs.
- Cache hot pastes in Redis to achieve sub-10ms read latencies.
- Implement dual expiration: lazy deletion on read and scheduled background S3 sweeping.
Topic Knowledge Assessment ๐ง
Step through 2 scenario questions to test your staff-level grasp.
Why should the raw 5MB text body of a paste be stored in AWS S3 rather than directly inside a relational PostgreSQL table column?
How clear and staff-actionable was this system breakdown?