How to open a large JSON file (100 MB to 1 GB) without crashing
· 7 min read
You download an export, a log dump or an API backup, double-click it, and your editor turns grey. A few minutes later it either shows the file or dies. This guide explains why that happens and walks through what actually works at 100 MB, 500 MB and 1 GB, with commands you can paste.
The short version: decide first whether you need to look at the file or extract something from it. Looking needs a viewer that doesn't build the whole thing at once. Extracting needs a streaming parser or jq.
Why editors choke on big JSON
Three separate limits get hit, usually in this order.
The parsed tree is much bigger than the file
A JSON file is compact text. Once parsed, every object, key and number becomes a separate allocation with its own overhead. I measured this on a 15.5 MB file of 200,000 small order objects (Intel i5-9300H, Linux):
| Tool | Peak memory | Time |
|---|---|---|
jq '.orders | length' |
231 MB | 0.6 s |
Python json.load |
142 MB | |
Node JSON.parse |
104 MB | |
jq --stream (filter shown below) |
3.6 MB | 6.0 s |
Python ijson.items |
14 MB | 0.4 s |
Plain jq used about 15 times the file size. Scale that to a 1 GB file and you need well over 10 GB of RAM for the same approach. Streaming parsers stay flat because they never hold more than one item.
Text editors try to highlight and fold everything
Editors also tokenize the text for colours, compute folding ranges and run a language server for validation. VS Code switches these off for big files: its source sets the threshold at 20 MB or 300,000 lines, and above 50 MB it stops syncing the file to extensions, so JSON validation and outline go away. Even with those optimizations, a few hundred megabytes of single-line JSON is painful to scroll, and a minified file with one giant line is the worst case for any text editor.
Browsers have a hard string limit
In Chrome, the largest string JavaScript can create is 536,870,888 characters (about 512 MB). Any web tool that does file.text() and then JSON.parse stops there, whatever your RAM. The usual error is "Invalid string length" or "Cannot create a string longer than 0x1fffffe8 characters". On top of that, parsing on the page's main thread freezes the tab for as long as it takes.
First: look at the file without opening it
Before picking a tool, find out what's inside. These commands read only the start of the file:
ls -lh export.json # how big is it really
head -c 2000 export.json # the first 2 KB: is it one array, an object, NDJSON?
If every line is its own JSON object, you don't have one big JSON document. You have NDJSON (JSON Lines), and everything gets easier: head, grep, split and wc -l work, and every tool can process it line by line. See viewing and querying NDJSON logs.
If it's one big document, note the top-level shape. Most large files are {"meta": ..., "items": [ ...millions of things... ]}. That path (items) is what you will stream over.
Option 1: jq, when the machine has the RAM
jq is the quickest way to answer a question about a file if it fits in memory. A few useful commands:
jq '.orders | length' big.json # how many items
jq -c '.orders[:2]' big.json # the first two items, compact
jq 'keys' big.json # top-level keys
jq -c '.orders[] | select(.status == "refunded") | .id' big.json > refunded.txt
The catch is memory. Use the table above as a rule of thumb: plan for 10 to 20 times the file size. For a 300 MB file on a 16 GB laptop that's fine. For 1 GB it probably isn't.
Option 2: jq --stream, slow but constant memory
--stream makes jq emit path/value events instead of building the document. Combined with fromstream and truncate_stream, you can rebuild one array element at a time:
jq -cn --stream '
fromstream(2 | truncate_stream(inputs | select(.[0][0] == "orders")))
| select(.status == "refunded")
| .id' big.json
The 2 is the depth to strip: path ["orders", 17, "status"] becomes ["status"], so each order is rebuilt as a standalone object. If the file is a bare top-level array, drop the select and strip one level: jq -cn --stream 'fromstream(1 | truncate_stream(inputs))' big.json.
On my 15.5 MB test file this ran in 6.0 s with 3.6 MB of memory, against 0.6 s and 231 MB for plain jq. Ten times slower, sixty times less memory. For a multi-gigabyte file on a small machine, that trade is the only one that works.
Option 3: a streaming parser in Python or Node
If the job is "go through every record and compute something", write ten lines of code with a streaming parser.
Python with ijson
pip install ijson
import ijson
from decimal import Decimal
refunded = 0
total = Decimal(0)
with open("big.json", "rb") as f:
for order in ijson.items(f, "orders.item"):
if order["status"] == "refunded":
refunded += 1
total += order["total"]
print(refunded, total)
"orders.item" means "each element of the orders array". For a file that is a bare array, use "item". ijson returns numbers with decimals as Decimal, which is why the total starts as one. It picks a fast C backend (yajl2_c) when it's available; on my test it was faster than json.load and used a tenth of the memory.
Node with stream-json
Node has no streaming JSON parser built in. The stream-json package fills the gap:
npm install stream-json
const fs = require('fs');
const { chain } = require('stream-chain');
const { parser } = require('stream-json');
const { pick } = require('stream-json/filters/Pick');
const { streamArray } = require('stream-json/streamers/StreamArray');
let refunded = 0;
chain([
fs.createReadStream('big.json'),
parser(),
pick({ filter: 'orders' }),
streamArray(),
])
.on('data', ({ value }) => { if (value.status === 'refunded') refunded++; })
.on('end', () => console.log(refunded));
(Tested with stream-json 1.9.1.) It's slower than ijson's C backend but works on files of any size. Don't try JSON.parse(fs.readFileSync(path, 'utf8')) on a 1 GB file: it fails at the string limit before parsing starts.
Option 4: load it into a database
If you will ask many different questions of the same file, import it once. DuckDB reads JSON and NDJSON directly with read_json_auto, and SQLite has JSON functions. You pay the import time once, then every query is fast and you get SQL. This is the right call for analysis work that goes on for days; for a one-off look it's overkill.
Option 5: split it into smaller files
When someone else needs the data in a tool with a size limit, split it. With jq streaming into NDJSON and the split command:
jq -cn --stream 'fromstream(2 | truncate_stream(inputs | select(.[0][0] == "orders")))' big.json > orders.ndjson
split -l 100000 orders.ndjson part_
Each part_aa, part_ab... holds 100,000 records, one per line. jq -s . part_aa > part_aa.json turns one back into a JSON array.
Option 6: a viewer that parses off the main thread
Sometimes you just need to look: scroll the structure, search for an ID, check what a record looks like at position 1.2 million. A viewer is the right tool for that, as long as it does three things: parses in a background thread (a Web Worker in a browser) so the window stays responsive, keeps the parsed data out of the UI and renders only the rows on screen, and reads the file in pieces instead of one string.
GigaJSON is a browser tool built around those rules. Open a file in the large JSON viewer or the full app and the parsing happens in a worker on your machine; nothing is uploaded. In our benchmark on the same i5-9300H laptop in Chrome, the tree appeared after 3.3 s for a 200 MB file, 10.1 s for 500 MB and 21.3 s for 1 GB (medians of three runs, October 2026), and the page stayed responsive the whole time. The page itself kept a heap of 5 to 7 MB at every size, because it only ever receives the rows on screen.
What you can do once it's open: browse the tree, see an array of objects as a table, search with Find, and run JSONPath or JMESPath queries. Some honest limits:
- The default largest-document setting is 512 MB. You raise it to 1 GB in Preferences.
- At 1 GB the worker needs a few GB of RAM, so on an 8 GB machine with other things open, close tabs first.
- Whole-document queries are slow at that size:
$..skuover the 1 GB file took about a minute. - jq in the browser is limited to about 20 MB of input (the WebAssembly build aborts past ~25 MB), so for big files use JSONPath/JMESPath there, or jq on the command line.
- Desktop viewers written in native code open 1 GB files much faster than any browser tool can. If you open giant files every day, that speed may matter more to you than not installing anything.
- The file is stored in the browser's own storage so you can come back to it, which takes roughly twice the file size on disk (the document plus its "Original" version).
Which option to pick
| You want to... | Use |
|---|---|
| Answer one question, file fits in RAM | jq |
| Process every record of a multi-GB file | ijson (Python) or stream-json (Node) |
| Same, without writing code, on a small machine | jq --stream |
| Ask many questions over days | DuckDB or SQLite |
| Hand the data to someone with a size limit | split into NDJSON parts |
| Browse, search and inspect up to 1 GB | a viewer that parses in a worker, such as the large JSON viewer |
One last tip: if you control how the file is produced, write NDJSON instead of one giant array. Every tool in this list handles it better, and so does grep.