How to Process a 2GB File on AWS: Architectural Patterns for S3 Streaming, Step Functions Distributed Map & ECS Fargate

A 2GB file sits at an awkward architectural inflection point on AWS. It is small enough that engineers often attempt to treat it like standard transactional data—attempting direct HTTP uploads, single-shot Lambda invocations, or reading the entire buffer into memory—yet large enough that every naive assumption fails with silent memory crashes, heap exhaustion, or 15-minute invocation timeouts. In this guide, I break down the concrete architectures for processing 2GB files across the AWS ecosystem: comparing single-instance streaming in AWS Lambda, massively parallel record chunking with Step Functions Distributed Map, sustained batch processing on ECS Fargate, and zero-compute columnar queries using Amazon Athena and DuckDB.

1. The Anatomy of the 2GB Constraint on AWS

To architect an effective file processing pipeline, you must first understand the hard physical and service boundary conditions imposed by AWS compute and networking primitives:

  • API Gateway & Load Balancer Ceilings: AWS API Gateway enforces a non-negotiable payload ceiling of 10MB on REST and HTTP APIs. Application Load Balancers (ALBs) enforce a 1MB default chunk limit. Any architecture that attempts to ingest a 2GB file through an HTTP POST endpoint on API Gateway will fail immediately with HTTP 413 (Payload Too Large). Ingestion must bypass the application layer entirely using Amazon S3 Presigned URLs with Multipart Upload.
  • AWS Lambda Memory & Ephemeral Storage: Lambda functions can be provisioned with up to 10,240MB (10GB) of RAM and up to 10,240MB of ephemeral disk in /tmp. However, allocating a 4GB or 6GB Lambda simply to buffer a 2GB file in memory is an expensive anti-pattern. Furthermore, the default ephemeral storage for Lambda is 512MB; writing a 2GB file to /tmp without explicitly raising the ephemeral storage configuration immediately triggers ENOSPC: no space left on device.
  • The 15-Minute Execution Timeout: AWS Lambda enforces a hard execution ceiling of 900 seconds (15 minutes). If a 2GB file contains 5,000,000 records and each record requires a 5ms external database write or validation call, processing the records sequentially takes over 6.9 hours. Single-threaded serverless compute cannot finish the job.
  • V8 Engine & Runtime Heap Limits: In Node.js 64-bit environments, the default V8 memory heap limit is approximately 1.4GB to 2GB. Attempting to execute fs.readFileSync() or calling await response.Body.transformToString() on a 2GB file in Node.js crashes the runtime with FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory. For deep-dive diagnostics on inspecting V8 garbage collection and memory profiling under heavy I/O pressure, see our guide on diagnosing Node.js memory leaks with heap snapshots and flamegraphs.

Navigating these limits requires matching your file data structure (uncompressed line-delimited records vs. nested binary archives) to the correct AWS compute pattern.

2. Architecture 1: Single-Instance S3 Streaming in AWS Lambda

When the 2GB file consists of sequential line-delimited data (such as CSV, NDJSON, or server logs) and the transformation is purely compute-bound (such as schema normalization, filtering, or hashing), you do not need multi-worker orchestration or multi-gigabyte RAM allocations. You can process the entire 2GB file inside a lightweight 512MB or 1,024MB Lambda function using Node.js stream pipelines with constant O(1) memory complexity.

The core principle is backpressure-aware pipelining: data is pulled from Amazon S3 in 64KB chunks, parsed row-by-row through a transform stream, and immediately pushed back to a destination S3 bucket via a streaming multipart upload. At no point does more than a few megabytes of data reside in memory.

Implementation: End-to-End Streaming with AWS SDK v3

Here is a complete, production-tested implementation using the Node.js stream/promises pipeline and the @aws-sdk/lib-storage multipart upload abstraction:

import { S3Client, GetObjectCommand } from '@aws-sdk/client-s3';
import { Upload } from '@aws-sdk/lib-storage';
import { pipeline } from 'node:stream/promises';
import { Transform, PassThrough, Readable } from 'node:stream';

const s3 = new S3Client({ region: 'us-east-1' });

export async function handler(event: { sourceBucket: string; sourceKey: string; targetBucket: string; targetKey: string }) {
  console.log(`Starting stream processing for s3://${event.sourceBucket}/${event.sourceKey}`);

  // 1. Fetch readable stream from S3 without buffering object to disk or RAM
  const getCommand = new GetObjectCommand({
    Bucket: event.sourceBucket,
    Key: event.sourceKey,
  });
  const response = await s3.send(getCommand);
  const s3Stream = response.Body as Readable;

  let rowCount = 0;
  let remainder = '';

  // 2. Line-by-line transform with strict backpressure and zero heap accumulation
  const lineFilterTransform = new Transform({
    transform(chunk: Buffer, _encoding, callback) {
      remainder += chunk.toString('utf8');
      const lines = remainder.split(/\r?\n/);
      remainder = lines.pop() ?? '';

      let output = '';
      for (const line of lines) {
        if (!line) continue;
        rowCount++;
        if (line.startsWith('TIMESTAMP') || line.includes('"status":"valid"')) {
          output += line + '\n';
        }
      }
      callback(null, output);
    },
    flush(callback) {
      if (remainder && (remainder.startsWith('TIMESTAMP') || remainder.includes('"status":"valid"'))) {
        rowCount++;
        callback(null, remainder + '\n');
      } else {
        callback();
      }
    }
  });

  // 3. PassThrough bridge into the multipart upload consumer
  const uploadStream = new PassThrough();

  const s3Upload = new Upload({
    client: s3,
    params: {
      Bucket: event.targetBucket,
      Key: event.targetKey,
      Body: uploadStream,
      ContentType: 'application/x-ndjson',
    },
    partSize: 10 * 1024 * 1024, // 10MB chunk parts (exceeds 5MB S3 minimum)
    queueSize: 4,
  });

  // Execute pipeline and multipart upload concurrently with fail-fast error propagation
  await Promise.all([
    pipeline(s3Stream, lineFilterTransform, uploadStream),
    s3Upload.done()
  ]);

  console.log(`Stream pipeline completed. Processed ${rowCount} rows with zero memory spikes.`);
  return { status: 'SUCCESS', processedRows: rowCount };
}

Performance & Cost Characteristics

On a Lambda provisioned with 1,769MB of RAM (which guarantees exactly 1 full vCPU thread), a 2GB CSV containing 12,000,000 rows streams and parses in approximately 42 seconds. Because memory usage remains flat at ~95MB throughout the execution, the cost of processing the entire 2GB file is approximately $0.0012.

When to use this pattern: Clean sequential transformations, redaction, compression, or format conversion where processing can complete comfortably within 10 minutes and requires zero network roundtrips per row. When processing high-throughput streaming workloads across tenant boundaries, pairing stream pipelines with multi-tenant domain dispatch and sharded data architectures ensures zero cross-tenant resource contention and deterministic memory ceilings.

3. Architecture 2: Massively Parallel Processing with Step Functions Distributed Map

Streaming inside a single Lambda function falls apart as soon as individual records require downstream network operations. If parsing each record requires querying DynamoDB, invoking an ML model endpoint, or verifying credentials against an external API, sequential execution will blow past the 15-minute execution limit.

The solution is AWS Step Functions Distributed Map. In Distributed Map mode, Step Functions acts as an autonomous distributed coordinator capable of spawning up to 10,000 parallel child workflow executions directly from an S3 dataset.

Configuration Parameter Recommended Value for 2GB File Architectural Rationale
ItemReader s3:getObject (CSV / JSON) Natively reads and slices the 2GB object in S3 without pre-splitting files or loading manifests into state memory.
ItemBatcher.MaxItemsPerBatch 1,000 to 5,000 records Groups records to balance invocation cold starts against single Lambda execution runtimes.
MaxConcurrency 50 to 200 workers Protects downstream relational databases (Amazon Aurora connection pools) or third-party APIs from request saturation.
ResultWriter s3:putObject (Export Bucket) Writes child task execution manifests and outputs directly to S3, bypassing the 256KB Step Functions execution output limit.

Amazon States Language (ASL) Specification

Here is the architectural state machine definition demonstrating how Distributed Map inspects an S3 object and orchestrates parallel worker tasks:

{
  "Comment": "Process 2GB S3 File via Distributed Map with Batching",
  "StartAt": "ProcessLargeS3File",
  "States": {
    "ProcessLargeS3File": {
      "Type": "Map",
      "ItemProcessor": {
        "ProcessorConfig": {
          "Mode": "DISTRIBUTED",
          "ExecutionType": "EXPRESS"
        },
        "StartAt": "ProcessRecordBatch",
        "States": {
          "ProcessRecordBatch": {
            "Type": "Task",
            "Resource": "arn:aws:states:::lambda:invoke",
            "OutputPath": "$.Payload",
            "Parameters": {
              "FunctionName": "arn:aws:lambda:us-east-1:123456789012:function:batch-record-worker",
              "Payload": {
                "records.$": "$.Items"
              }
            },
            "End": true
          }
        }
      },
      "ItemReader": {
        "Resource": "arn:aws:states:::s3:getObject",
        "ReaderConfig": {
          "InputType": "CSV",
          "CSVHeaderLocation": "FIRST_ROW"
        },
        "Parameters": {
          "Bucket.$": "$.SourceBucket",
          "Key.$": "$.SourceKey"
        }
      },
      "ItemBatcher": {
        "MaxItemsPerBatch": 2500,
        "BatchInput": {
          "jobId.$": "$.JobId"
        }
      },
      "MaxConcurrency": 80,
      "ResultWriter": {
        "Resource": "arn:aws:states:::s3:putObject",
        "Parameters": {
          "Bucket.$": "$.OutputBucket",
          "Prefix.$": "States.Format('{}/results', $.JobId)"
        }
      },
      "End": true
    }
  }
}

Fault Isolation & Partial Retries

The primary advantage of Distributed Map over a single monolith is blast-radius containment. If 15 records in a 5,000,000-record file contain invalid characters or encounter a transient downstream network failure, Step Functions records those specific failed items in the S3 result manifest. The remaining 4,999,985 records complete successfully. You can retry the failed subset without reprocessing the entire 2GB file from scratch.

4. Architecture 3: Sustained Batch Processing on AWS ECS Fargate

While serverless streaming and Distributed Map are ideal for structured datasets, certain workloads cannot be broken into independent rows or executed within Lambda constraints:

  • Monolithic Binary CLI Tools: Transcoding a 2GB 4K video using ffmpeg, processing a 2GB GeoTIFF raster using GDAL, or decompressing an encrypted multi-file 7-zip archive.
  • Multi-Hour Computational Runtimes: Machine learning inference, graph traversal, or 3D rendering jobs that require 30 to 90 minutes of continuous compute.
  • Large Scratch Disk Requirements: When algorithms require random multi-gigabyte disk seeks that overwhelm Lambda ephemeral storage.

In these scenarios, the recommended architecture is an asynchronous, event-driven AWS ECS Fargate Task orchestrated via Amazon SQS.

Event-Driven Fargate Architecture Flow

1. Client uploads 2GB file to S3 via Presigned Multipart URL.
2. S3 emits an ObjectCreated notification event to Amazon EventBridge.
3. EventBridge enqueues the task payload into an Amazon SQS FIFO queue.
4. An ECS worker service (or AWS Batch compute environment) polls SQS and triggers a container task.
5. The task allocates up to 200GB of ephemeral storage, streams or mounts the S3 object, processes the workload, and writes the output back to S3.
6. Upon completion, the worker deletes the SQS message and terminates cleanly.

Fargate Ephemeral Storage vs. Amazon EFS

ECS Fargate tasks provide up to 200GB of configurable ephemeral storage (backed by NVMe SSDs). For a 2GB file processing job, configuring 20GB of ephemeral storage allows downloading the source file, unpacking intermediate staging buffers, and generating the destination artifact entirely on local container disk:

{
  "family": "file-processor-task",
  "cpu": "2048",
  "memory": "4096",
  "networkMode": "awsvpc",
  "requiresCompatibilities": ["FARGATE"],
  "ephemeralStorage": {
    "sizeInGiB": 25
  },
  "containerDefinitions": [
    {
      "name": "processor",
      "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/media-processor:latest",
      "essential": true,
      "environment": [
        { "name": "TEMP_DIR", "value": "/tmp/workspace" }
      ]
    }
  ]
}

By leveraging Fargate Spot instances, you receive up to a 70% discount compared to standard on-demand pricing. For long-running, fault-tolerant batch transformations, Fargate Spot is substantially more economical than provisioning multiple high-memory Lambda functions.

5. Architecture 4: Zero-Compute Columnar Analysis with Athena & DuckDB

Before building a complex compute pipeline, evaluate whether you actually need to transform the file. In many enterprise systems, files are uploaded simply so that analysts or reporting dashboards can query specific metrics (e.g. "filter transactions where region is West and calculate total spend").

Writing custom Node.js or Python code to parse a 2GB file just to execute aggregations is architectural waste. AWS provides zero-compute serverless query engines that can query 2GB files directly in place.

Approach A: Amazon Athena & Columnar Parquet

If you convert the 2GB raw CSV file to Apache Parquet with Snappy or ZSTD compression, the file size drops from 2,000MB to roughly 180MB–250MB. Because Parquet is a columnar format, Amazon Athena (built on Trino/Presto) does not scan the entire file. When you query two columns out of fifty, Athena uses S3 HTTP range requests to download only the byte offsets containing those two columns.

SELECT 
  customer_id, 
  COUNT(*) AS transaction_count,
  SUM(amount) AS total_revenue
FROM s3_financial_data.settlements_parquet
WHERE settlement_date >= '2026-01-01'
GROUP BY customer_id
HAVING SUM(amount) > 5000;

A query against a 2GB Parquet dataset typically executes in 1.4 seconds, scans less than 35MB of data, and costs less than $0.0002 on Athena's $5.00 per TB pricing model.

Approach B: In-Process DuckDB on AWS Lambda

For sub-second analytical queries without provisioning an Athena workgroup, you can embed DuckDB directly inside an AWS Lambda function. DuckDB's httpfs extension enables native HTTP range requests against S3 objects:

import duckdb

def handler(event, context):
    con = duckdb.connect(':memory:')
    # Load HTTPFS and AWS credential resolver for Lambda IAM execution role
    con.execute("INSTALL httpfs; LOAD httpfs;")
    con.execute("INSTALL aws; LOAD aws; CALL load_aws_credentials();")
    con.execute("SET s3_region='us-east-1';")
    
    # Query S3 Parquet file directly via HTTP range requests
    query = """
        SELECT category, SUM(price) as total_sales
        FROM read_parquet('s3://my-data-bucket/exports/2gb_inventory.parquet')
        GROUP BY category
        ORDER BY total_sales DESC
        LIMIT 10;
    """
    df = con.execute(query).df()
    return df.to_dict(orient='records')

The entire query executes inside a 512MB Lambda in under 800ms, transferring only the requested byte ranges across the network without downloading the remaining 95% of the file.

6. Architectural Decision Matrix

Use the following comparison matrix to select the optimal AWS pattern based on file format, latency requirements, and computational complexity:

Architectural Pattern Optimal Data Format Memory Footprint Execution Limit Cost Profile Primary Advantage
Lambda S3 Stream CSV, NDJSON, Plaintext ~95MB to 128MB (Flat O(1)) 15 Minutes Lowest ($0.001 per run) Zero infrastructure, instantaneous startup, O(1) RAM.
Step Functions Dist. Map CSV, JSON Array, S3 Manifest 128MB per child worker Up to 1 Year (Workflow) Moderate (State transitions + Lambda) Automatic row chunking, 10k parallelism, granular per-item retries.
ECS Fargate Batch Video, Audio, Zip Archives, Geospatial Up to 120GB RAM No limit (Hours / Days) Spot discounts (70% off) POSIX filesystem, custom native CLI binaries (FFmpeg, GDAL).
Amazon Athena / DuckDB Parquet, ORC, Compressed CSV Zero server memory 30 Minutes (Athena) Pay-per-scanned-byte ($5/TB) Zero ETL pipeline code, direct ad-hoc SQL over raw S3 files.

7. Critical Operational Traps & Production Hardening

Regardless of which architecture you select, several subtle edge cases can derail a 2GB processing pipeline in production:

  • S3 Multipart Upload Minimum Part Size: When streaming output back to S3 using the AWS SDK, the minimum part size for multipart uploads is exactly 5MB (5,242,880 bytes), with the exception of the final part. If your stream flushes chunks smaller than 5MB to S3's UploadPart API, S3 throws EntityTooSmall: Your proposed upload is smaller than the minimum allowed size. Always configure partSize in @aws-sdk/lib-storage to at least 5MB (10MB is optimal).
  • Downstream Database Saturation: Streaming 2GB of data into PostgreSQL or MySQL row-by-row using individual INSERT statements will exhaust connection pools, trigger write lock contention, and fail. When writing to relational databases from a stream, buffer records into micro-batches of 1,000 and execute bulk copy commands (e.g. pg-copy-streams for Postgres or LOAD DATA LOCAL INFILE for MySQL).
  • Warm Container Memory Creep: In AWS Lambda, global objects, unclosed readable streams, and cached arrays persist across execution contexts. If your handler attaches event listeners to global event emitters without unregistering them, memory usage will creep upwards on successive invocations until the container crashes. Always ensure streams are piped via stream/promises.pipeline, which guarantees deterministic teardown on completion or error.
  • Idempotency Under S3 Event Retries: Amazon S3 event notifications guarantee at-least-once delivery. If network hiccups occur, S3 may emit duplicate ObjectCreated events for the same 2GB file. Use the S3 object's ETag or S3 Object Version ID to create a conditional lock entry in DynamoDB before kicking off processing. Designing idempotent pipelines with deterministic retry containment mirrors the non-destructive recovery principles detailed in our lossless recovery and architecture hardening guide.

Core Architectural Decision Rule

When processing a 2GB file on AWS: if the data is sequential and linear, stream it in Lambda. If records require external API calls or isolated failure recovery, chunk it with Step Functions Distributed Map. If processing requires heavy CLI binaries or runs longer than 15 minutes, dispatch it to ECS Fargate. And if you only need answers to analytical questions, query it directly with Athena or DuckDB.