JSON Lines and Streaming Large JSON Files

Advanced
13 min

JSON Lines and Streaming Large JSON Files

JSON.parse is fast, but the whole document must be in memory as one string before parsing starts. That is fine for an API response and hopeless for a 4 GB export. In this lesson you will learn the JSON Lines format that makes large datasets streamable and appendable, how to read and write it in Node.js and Python, how to convert it with jq, and how to stream-parse a big conventional JSON file when you do not control its format.

Why One Big JSON Document Is a Problem

A JSON array of a million records has three practical limits. Memory: the text must be loaded completely, V8 caps a single string at roughly 512 MB, and the parsed tree occupies several times the text size. Latency: nothing can be processed until the final ] arrives. Mutability: appending one record means rewriting the entire file. Logs, exports and event streams hit all three.

The JSON Lines Format

JSON Lines (also called NDJSON, newline-delimited JSON) fixes this with one rule: one complete JSON value per line, separated by \n, with no surrounding array and no commas between lines. Files use the .jsonl or .ndjson extension and are usually served as application/x-ndjson.

json
{"ts":"2026-09-27T10:00:01Z","level":"info","msg":"server started","port":3000} {"ts":"2026-09-27T10:00:07Z","level":"warn","msg":"slow query","ms":842} {"ts":"2026-09-27T10:00:09Z","level":"error","msg":"db timeout","retry":true}

Each line is independently parseable: a record is appended with a single write, a reader can process line one before line two exists, grep, head, wc -l and split work directly, and a corrupted line spoils only itself. This is why it is the format of application and container logs, BigQuery and Elasticsearch bulk data, and machine-learning training sets.

Reading and Writing JSON Lines

In Node.js, readline over a file stream yields one line at a time with constant memory, as the example at the top of the lesson shows. Writing is JSON.stringify(record) + "\n" appended to the file; the record must be minified so it stays on one line. Python's file object already iterates by line:

python
import json errors = 0 with open("events.jsonl", encoding="utf-8") as f: for line in f: if not line.strip(): continue event = json.loads(line) if event["level"] == "error": errors += 1 with open("events.jsonl", "a", encoding="utf-8") as f: f.write(json.dumps({"level": "info", "msg": "done"}) + "\n")

pandas reads the format directly with pd.read_json("events.jsonl", lines=True).

Converting With jq

jq treats a stream of JSON values as naturally as a single document, so it converts between the two forms in one line and filters JSON Lines record by record:

bash
jq -c '.[]' big-array.json > records.jsonl # array -> JSON Lines (compact, one per line) jq -s '.' records.jsonl > big-array.json # JSON Lines -> array (-s slurps all inputs) jq -c 'select(.level == "error")' records.jsonl # filter, keeping the line format # True streaming for an array too large to load: emits each top-level element jq -cn --stream 'fromstream(1 | truncate_stream(inputs))' big-array.json > records.jsonl

The first command still loads big-array.json fully; only the --stream form reads it incrementally.

Streaming Parsers for Ordinary JSON

Sometimes the big file is a plain JSON array you did not produce. A streaming parser emits elements as they are parsed instead of building the whole tree. In Node.js the stream-json package is the standard choice:

javascript
const { createReadStream } = require("node:fs"); const { parser } = require("stream-json"); const { streamArray } = require("stream-json/streamers/StreamArray"); let active = 0; createReadStream("big-array.json") .pipe(parser()) .pipe(streamArray()) .on("data", ({ value }) => { if (value.active) active++; }) .on("end", () => console.log({ active }));

Memory stays flat regardless of file size because each element is released after the handler runs. Python's ijson (ijson.items(f, "item")) and Go's json.Decoder (dec.Decode(&v) while dec.More()) offer the same pattern. Over HTTP, servers stream JSON Lines by flushing one line per record, and browsers read response.body with a TextDecoder, splitting on newlines.

Common Mistakes

  • Pretty-printing records in a .jsonl file. A record spanning several lines breaks every line-based reader. Always write minified records.
  • Adding commas or an outer array. That turns the file back into one big document.
  • Ignoring the last line. Writers should always end with \n; readers should tolerate a missing final newline and blank lines.
Quick Quiz
Question 1 of 3

What distinguishes a JSON Lines file from a JSON array?

Key Takeaways

  • JSON.parse needs the entire document in memory; large arrays hit memory, latency and append limits.
  • JSON Lines stores one minified JSON value per line, which makes data appendable, greppable and streamable.
  • Read .jsonl with readline in Node.js or by iterating the file in Python; write with JSON.stringify(record) + "\n".
  • jq -c '.[]' and jq -s '.' convert between arrays and JSON Lines; --stream handles arrays too big to load.
  • For huge conventional JSON, use a streaming parser such as stream-json, ijson or Go's json.Decoder.

Next lesson: JSON in Databases: PostgreSQL JSONB and MongoDB — store, query and index JSON documents inside relational and document databases.

JSON Lines and Streaming Large JSON Files - JSON | CodeYourCraft | CodeYourCraft