Solutioning Lab · Working Session Tool

MrBeast Media Platform Solutioning Lab

IBM + MrBeast working sessionOptions, decision points, bake-offs, roadmap, discovery gaps.
MrBeast
IBM · MrBeast — joint working architecture

A future-state global media data and AI architecture, designed in the open.

This is a solutioning lab for understanding, comparing, and refining options for a rapidly growing on-premises video estate — storage economics, global data movement, lifecycle management, and AI-driven discoverability of the archive.

ConfirmedRecommendedEvaluateAssumptionOpen decisionLabels used throughout to separate fact from working view.
Working architecture for joint discussion — subject to validation with MrBeast engineering and production teams.
01Section

Executive Overview

What this workspace is, the challenge it addresses, and the target state proposed for joint discussion.

Working architecture for joint discussion — subject to validation with MrBeast engineering and production teams. This is a solutioning lab, not a deployable system: no live integrations, no media ingest, no telemetry, and no operation of any MrBeast infrastructure.
The business challenge

Five forces are converging at once: storage economics at 44 PB and climbing, data gravity that makes cloud egress prohibitive, international movement between US and UK teams, time-to-content as the binding production constraint, and archive intelligence that turns retained footage back into usable creative material.

Design philosophy

Store every asset on the appropriate economic tier. Move only the data that needs to move. Keep intelligence about every asset immediately available.

Target-state statement
Recommended

An on-premises-first, IBM-led hybrid architecture that preserves the existing NetApp investment while adding three things it does not have today: lifecycle economics across active, object, and archive tiers; global data movement between US, UK, and future international sites; and an AI intelligence layer that keeps every asset findable regardless of where the master physically sits.

Lifecycle economics
Active → object → archive by policy
Global movement
One namespace, accelerated transport
Archive intelligence
Metadata online even when masters are cold
~44 PB
Total current estate
Sum of stated tiers
Confirmed
12 PB
Tier 1 — active production
As shared by MrBeast team
Confirmed
15 PB
Tier 2 — near-line
As shared by MrBeast team
Confirmed
17 PB
Tier 3 — retained / archive
As shared by MrBeast team
Confirmed
~300 TB
Peak ingest per day
Heavy production periods
Confirmed
8K
Capture transition underway
Drives per-project growth
Confirmed
~12
Infrastructure team size
Lean operating model
Confirmed
Reading this in Joint Solution view

Emphasis on options, decision points, the Hammerspace vs IBM Storage Scale bake-off, roadmap sequencing, and the discovery gaps both teams need to close together.

02Section

Business Problems

Nine pressures shaping the architecture, stated in the terms each audience cares about.

Rapid storage growth

Capacity is growing faster than a lean team can procure, rack, and operate it.

Growth compounds cost and management effort at the same time. Without tiering, every new petabyte is priced like production storage.

Confirmed: 12 / 15 / 17 PB tiers.

Data gravity and cloud egress economics

At this scale, moving data to and out of public cloud can cost more than storing it.

Cloud is not ruled out, but at 44 PB and rising, egress and retrieval charges dominate the model.

Working assumption: on-premises-first for masters.

International transfer and collaboration

US and UK teams need working access to the same material without waiting on full copies.

Global production speed depends on how quickly a UK editor can start work, not on how fast a full transfer completes.

Confirmed: UK workflows are in scope.

Shipping-drive risk and delay

Physical media moves add days, logistics load, and integrity exposure.

Drives in transit are unbilled time and unmanaged risk — loss, damage, or silent corruption.

Confirmed pain point.

Time-to-content

Time is the binding constraint across capture, ingest, edit, and delivery.

Every hour saved between capture and first edit is directly recoverable production capacity.

Stated by the MrBeast team as the biggest enemy.

Archive discoverability

Historic footage is retained but hard to find at the moment of creative need.

Archive value only materializes if a producer can find the right shot in seconds instead of days.

Objective stated by MrBeast leadership.

8K growth

Higher resolution multiplies bytes per shooting day across every downstream tier.

Format change amplifies every other problem on this page simultaneously.

Confirmed direction of travel.

Lean operations team

~12 people support a growing global estate.

Any architecture that adds operational headcount is the wrong architecture.

Confirmed team size.

Future multi-site expansion

Additional production and post locations are expected.

The design should make the second and third site an incremental step, not a re-architecture.

Working assumption — site count to be confirmed.
03Section

Current State

How content moves today, with confirmed facts separated from working assumptions.

Current-state flow
ConfirmedAssumption
01
On-set / 8K captureConfirmed

Camera cards, up to ~300 TB/day at peak

02
High-speed ingestAssumption

Offload, checksum, and hand-off to production storage

03
Tier 1 — 12 PBConfirmed

Active editorial performance storage (NetApp)

04
Tier 2 — 15 PBConfirmed

Near-line / recently completed work

05
Tier 3 — 17 PBConfirmed

Retained content and longer-term holdings

06
Local & international transferConfirmed

Network transfer plus physical drive shipment for some UK movement

Tier boundaries above reflect capacity as shared by the MrBeast team. The rules that currently move data between tiers, and the split between unique content and protection copies, are still to be confirmed.
Discovery notes
  • ConfirmedUK post-production house involved in workflows
  • ConfirmedDocumentary production alongside episodic content
  • ConfirmedExisting NetApp footprint and support relationship
  • ConfirmedPeak ingest up to ~300 TB/day
  • ConfirmedTime described as the biggest operational enemy
  • ConfirmedCloud viewed as potentially too expensive at this scale
  • ConfirmedPhysical drive shipping used for some international movement
  • AssumptionTier boundaries are capacity-based; policy thresholds not yet documented
  • AssumptionUnique content vs protection copies inside 44 PB not yet separated
  • AssumptionMAM/DAM and NLE stack details still to be gathered
  • AssumptionUS↔UK circuit bandwidth and latency not yet measured
  • AssumptionProxy generation point in the workflow to be confirmed
04Section

Target Architecture

A conceptual future-state model. Select any component to review its role, value, tradeoffs, and alternatives.

Working architecture
Representational diagram for discussion. Nothing here connects to live systems — select any component to open its role, benefits, tradeoffs, dependencies, and alternatives.
Capture
On-set / 8K acquisition
Active production storage
Editorial performance tier
Global data layer
Primary architectural bake-off
Sites
One namespace, many locations
Data movement
Site-to-site and tier-to-tier
Storage lifecycle
Right asset, right economic tier
Data & intelligence plane
Intelligence stays online regardless of tier
05Section

Technology Components

Every candidate component with role, benefits, pros, cons, dependencies, alternatives, and recommendation status.

Recommendation status reflects the current working view of the joint architecture team and is open to challenge.

06Section

Storage Lifecycle

Masters move down the tiers; intelligence about every asset stays online.

Thresholds below are editable illustrative assumptions for discussion — not recommendations. They should be tuned against measured access patterns during the storage economics assessment.
Lifecycle flow
STAGE 1
New content
Ingest landing

Camera cards and on-set capture land at high speed; checksums and proxies generated as early as workflow allows.

Stays online
MasterProxy (generated)Technical metadata
STAGE 2
Active production
NetApp / Storage Scale

Full editorial performance while the project is being cut, reviewed, and delivered.

Stays online
MasterProxyTranscriptThumbnailsEmbeddings
STAGE 3
Completed project
Policy evaluation

Delivery milestone triggers lifecycle evaluation rather than a manual cleanup task.

Stays online
ProxyTranscriptThumbnailsEmbeddings
STAGE 4
IBM COS (on-prem)
Warm / cool object

Master moves to object economics but remains immediately readable for re-cuts and repurposing.

Stays online
ProxyTranscriptThumbnailsEmbeddings
STAGE 5
IBM Deep Archive
Cold — Deep Archive + Diamondback

Master preserved at the lowest sustainable cost; recall is a deliberate, tracked action.

Stays online
ProxyTranscriptThumbnailsEmbeddings

The master moves down the tiers. The intelligence about the master — proxy, transcript, thumbnails, embeddings, technical metadata — stays online throughout, so discovery never depends on where the bytes live.

Policy thresholds

Editable assumptions

Assumption
Move from active to object after project delivery +
days
Move from object to deep archive after last access +
days
Keep proxies, transcripts and embeddings online for
years
Retain masters for a minimum of
years
Conceptual UX mockup

“Find the moment the doors open in the warehouse reveal”

Mockup only

Illustrative only — not a live search system and not connected to any archive. It shows how one result set can span tiers while every result remains discoverable.

Ep. 212 — warehouse reveal, wide push-inActive

"the moment the doors open" — transcript + scene match

NetApp Tier 1Immediate
Documentary B-roll — coastal drone, golden hourWarm

visual embedding match, no dialogue

IBM COS (on-prem)Seconds
Ep. 088 — original challenge setupCold

OCR on on-screen graphic + transcript

IBM Deep ArchiveRecall request — illustrative SLA to be defined
UK unit — studio interview, second cameraWarm

speaker + scene detection

IBM COS (on-prem), UK siteSeconds
07Section

Scenario Modeling

A client-side planning calculator. Change any assumption to see directional implications immediately.

Every input below is an editable planning assumption, not a quote. All outputs are directional until validated against actual pricing, utilization, replication factors, and access-pattern data.
Inputs

Editable assumptions

Estate
Distribution
Sites & network
Illustrative cost index

The cost index is a unitless relative scale (active = 100 by default) used only to show directional differences between tiers. It is not pricing.

Total today
44 PB
Projected year 3
108.3 PB
Projected year 5
197.3 PB

Projected capacity by tier

Illustrative

Directional tier economics

Directional

Year-5 estate priced entirely on the active tier vs the modelled tier distribution, using the relative cost index above.

All-active (no lifecycle)19,730 idx
Tiered lifecycle model6,363 idx
Premium storage avoided (directional)
68%

Relative index only — not a cost saving estimate.

US↔UK transfer scenarios

Raw bandwidth math

Time to move a given volume at 10 Gbps with 70% usable efficiency. Excludes contention, protocol behaviour, and endpoint storage throughput.

Single shoot day (10 TB)3.2 hrs
Typical project (50 TB)15.9 hrs
Large project (250 TB)3.3 days
Peak ingest day4 days
Full Tier 1 estate158.7 days

Planning implication: a global namespace plus proxy-first workflows usually beats moving full masters — the transfer that matters is the one you avoid.

Site & retention context

Modelled across 2 data centre location(s) with a 10-year retention assumption. Capacity shown is logical unique content; replication and protection copies are excluded until the storage economics assessment separates them.

08Section

Architecture Options

Three viable models compared against neutral decision criteria.

All three models are viable. Criteria are scored on a neutral 1–5 scale to support discussion, not to rank vendors. Scores are a working view and expected to change as discovery evidence arrives.

Option A — IBM-centric

Single-vendor coherence

IBM Storage Scale as the global data layer, IBM COS and Deep Archive for lifecycle, IBM data and AI plane throughout.

Simplicity
Operational burden (lean team)
M&E specialization
Flexibility
IBM integration
Preserves existing NetApp
Strengths
  • One support and escalation path across storage, movement, and AI
  • Tightest integration between tiers and the intelligence plane
  • Simplest commercial and lifecycle alignment
Tradeoffs
  • Less media-industry-specific tooling at the namespace layer
  • Reduced flexibility to adopt best-in-class point solutions later

Option B — Best-of-breed

Specialist per layer

Hammerspace for the global data layer, NetApp retained for active production, mixed AI services selected per capability.

Simplicity
Operational burden (lean team)
M&E specialization
Flexibility
IBM integration
Preserves existing NetApp
Strengths
  • Strong media-and-entertainment fit at the namespace and workflow layer
  • Maximum freedom to swap components as the market moves
  • Preserves the existing NetApp investment fully
Tradeoffs
  • More vendors to integrate, support, and upgrade with ~12 people
  • Integration and accountability sit with the customer unless contracted otherwise

Option C — Recommended hybrid

Working recommendation
Working recommendation

Keep NetApp for active production, run the Hammerspace vs IBM Storage Scale bake-off for the global layer, standardise on Aspera, IBM COS, Deep Archive + Diamondback, and the watsonx intelligence plane.

Simplicity
Operational burden (lean team)
M&E specialization
Flexibility
IBM integration
Preserves existing NetApp
Strengths
  • No disruption to production storage on day one
  • Economics improve immediately via object and archive tiers
  • Global-layer choice is decided by evidence, not by architecture default
Tradeoffs
  • Requires the bake-off to be run and concluded
  • Two-vendor storage estate persists at least through the transition
09Section

Decision Support

The decisions that need to be made together, with recommendation, rationale, and the evidence required.

Owner and status fields are editable locally in this browser session for working-session use. Nothing is stored or transmitted.

Global data layer: Hammerspace vs IBM Storage Scale

Open decision
Recommendation
Run a structured bake-off; no default winner.
Rationale
Both can deliver a global namespace. The differentiators are M&E workflow fit, assimilation of the existing NetApp estate, and operability for a 12-person team.
Evidence needed
  • US↔UK latency and bandwidth measurements
  • Representative project transfer and open-in-NLE tests
  • Operability walk-through with the infrastructure team
Owner
Status

Future Tier-1 expansion approach

Open decision
Recommendation
Expand NetApp near-term; evaluate Storage Scale for 8K/AI-intensive growth.
Rationale
Avoids disruption now while keeping a higher-throughput option open for the 8K curve.
Evidence needed
  • 8K per-project footprint
  • Tier-1 utilization and headroom
  • AI pipeline concurrency targets
Owner
Status

Hot / warm / cold policy thresholds

Open decision
Recommendation
Start with illustrative thresholds and tune against measured access data.
Rationale
Thresholds set without access-pattern data will either strand cost or frustrate editors.
Evidence needed
  • Access frequency at 7 / 30 / 90 / 180 / 365 days
  • Re-cut and repurposing frequency
Owner
Status

Proxy strategy for AI processing

Open decision
Recommendation
Process AI against proxies, not masters; confirm where proxies are generated.
Rationale
Proxy-based processing keeps AI cost and data movement bounded and avoids archive recall for indexing.
Evidence needed
  • Current proxy format and generation point
  • Proxy quality sufficiency for OCR/VLM
Owner
Status

Video AI services and model selection

Open decision
Recommendation
Evaluate 2–3 stacks on real footage before committing.
Rationale
Accuracy on this specific content style, not benchmark scores, should decide it.
Evidence needed
  • Labelled evaluation clips
  • Accuracy and cost per hour of footage
  • Language/accent coverage
Owner
Status

Multi-site topology

Open decision
Recommendation
Design a repeatable site pattern; confirm number and location of future sites.
Rationale
A repeatable pattern makes each new site an incremental deployment rather than a redesign.
Evidence needed
  • Confirmed future site list
  • Per-site production profile
  • Local vs central archive preference
Owner
Status

WAN sizing and Aspera validation

Open decision
Recommendation
Measure current circuits, then size for target time-to-first-edit.
Rationale
Transfer acceleration is bounded by the underlying circuit; sizing must be evidence-based.
Evidence needed
  • Circuit inventory
  • Aspera test transfers on real project sizes
  • Peak concurrency profile
Owner
Status
10Section

POC / Roadmap

Three proposed validation workstreams that would replace assumptions with measured evidence.

Proposed validation workstreams for joint consideration. None of these are approved, scheduled, or contracted — scope, sequence, and participation would be agreed together.
Workstream 1

Storage Economics Assessment

Objective
Establish a validated baseline of what is stored, how it is accessed, and what each tier actually costs to serve.
Example activities
  • Inventory NetApp products, raw vs usable capacity, and protection copies
  • Separate unique content from replication and backup copies
  • Collect access-pattern data at 7 / 30 / 90 / 180 / 365 days
  • Model tier distribution against candidate lifecycle policies
Success criteria
  • Agreed unique-content baseline
  • Access-pattern curve accepted by both teams
  • Directional tier-distribution model both teams believe
Inputs required from MrBeast
  • Storage inventory exports
  • Utilization and access telemetry
  • Retention requirements
IBM participation
Storage architects and economics modelling support
Expected decision outcome

Decision input for lifecycle thresholds and the Tier-1 expansion approach.

Workstream 2

Global Data Movement & Access Validation

Objective
Prove that US↔UK collaboration can happen over the network at production speed, without shipping drives.
Example activities
  • Measure existing circuit bandwidth, latency, and loss
  • Run Aspera transfers using representative project sizes
  • Bake-off Hammerspace and IBM Storage Scale on the same test workflow
  • Time an editor opening a remote project end to end
Success criteria
  • Measured time-to-first-edit for a remote project
  • Transfer integrity verified across all test runs
  • Bake-off scorecard completed against agreed criteria
Inputs required from MrBeast
  • Network access and test windows
  • Representative project data set
  • Editor participation
IBM participation
Aspera and global data layer specialists; joint test facilitation
Expected decision outcome

Global data layer decision and WAN sizing recommendation.

Workstream 3

Content Intelligence Pilot

Objective
Demonstrate natural-language discovery over a bounded slice of the archive using real footage.
Example activities
  • Select a representative corpus (for example, one season plus documentary B-roll)
  • Run ASR, OCR, scene detection, and visual embedding over proxies
  • Normalise and index through unstructured.io into watsonx.data
  • Test editorial search use cases in watsonx.ai with real producers
Success criteria
  • Agreed accuracy threshold on the evaluation set
  • Producers find target shots faster than the current method
  • Cost per hour of processed footage understood
Inputs required from MrBeast
  • Pilot corpus and proxies
  • Existing transcripts/metadata if available
  • Producer/editor time
IBM participation
watsonx and data engineering resources; pilot design and evaluation
Expected decision outcome

Model selection decision and a scoped plan for archive-wide indexing.

11Section

Open Questions

The discovery checklist both teams can work through together.

Collaborative discovery checklist. Checkboxes and notes are held in this browser session only — nothing is saved, sent, or shared automatically.
0 / 12
Estate
Access
Workflow
Network
Growth
Archive
AI
Governance