Skip to content

Large files

A file larger than one request is sent in parts. The deployment checks every part as it arrives, checks the whole file once all parts are there, and saves nothing to the project until you confirm. Up to the deployment’s request limit nothing changes: a file goes in one request, as before.

Two numbers decide what happens, and the deployment states both (Python: client.served_limits()):

LimitWhat it boundsPackaged exampleWith nothing configured
upload_bytesone request, and so one part32 MiB8 MiB
file_bytesthe largest original4 GiBequal to upload_bytes

Seismic volumes (SEG-Y) can be as large as file_bytes. Other files sent in parts are read whole by their reader, so they stay within 64 MiB. What is served stays bounded as before: one slice or representation is at most 32 MiB.

  1. Choose the file on the upload page. The page says “This deployment accepts files up to 4 GiB.” and, for a file larger than one request, “The file is sent to this deployment and checked there. Nothing is saved to the project until you confirm.”
  2. Press Upload and check this file (2.6 GiB). The page first reads the file on your computer (“Reading your file on this computer: 38%”), then sends it: “Sending 1.2 of 2.6 GiB (47%). You can pause, or close this page and continue later.” Pause stops after the current part; Cancel upload discards what was sent.
  3. “Checking your volume on the server. This can take a few minutes; you can leave this page.” The check reads every trace header and stores each inline as a verified chunk.
  4. Review what the deployment found (for a volume: “Inlines 1595-2393, crosslines 1142-1620”), tick that you have permission to retain the file, and press Save as private file. Discard removes it instead.

Continue later. Closing the page keeps the unfinished upload in this browser for this project. Opening the upload page again says “An upload of NAME (47%) was not finished. Choose the same file to continue.” Choosing it reads the parts the deployment already holds on your computer again and compares them (“Checking that this is the same file…”); only the missing parts are sent. Another file is refused: “This is not the file you started with. Choose the same file, or discard the unfinished upload.” An unfinished upload is removed after 7 days.

from ophiolite import Client
client = Client.from_configuration("configuration.json")
print(client.served_limits()["file_bytes"]) # 4294967296 in the packaged example
volume = client.upload_data(
"Q13a_AMSTELFIELD_4501NEP_KPSDM_FINAL_FULL_STACK_TIME.SEGY", profile="segy/1",
declared={"z_domain": "time"}, name="Q13a full stack (time)",
attribution="Source: NLOG.NL ...", audience=[], rights_confirmed=True,
command_id="q13a-full-stack-time") # the same command continues an interrupted upload
inline = client.read_slice(volume.asset_id, volume.revision, "inline", 1595)
client.download_original(volume.asset_id, volume.revision, "copy.segy") # by ranges, verified

upload_data reads the file one part at a time, so memory stays one part whatever the file’s size. Calling it again with the same command_id after an interruption sends only the parts the deployment does not hold yet. The revision is the file’s SHA-256; the SDK refuses a result with any other. download_original reads by byte ranges and leaves nothing at the destination unless the whole file matches the digest the deployment states.

The notebook Send a large file does the same with a synthetic volume.

You seeWhyCode
“This file is 5.2 GiB. This deployment accepts files up to 4 GiB. Ask your administrator to raise the limit, or crop the volume.”over file_bytes; nothing was sentfile-too-large (413)
“This deployment has 1.5 GiB of room left for uploads. This file needs 2.6 GiB.”over a deployment total or the free diskno-room (413)
“Part 3 arrived damaged twice. Check your connection and continue.”a part did not match its digest twicepart-corrupt (422)
“The copy held by this deployment does not match your file. Nothing was saved; upload it again.”the whole file did not matchfile-mismatch (409)
“Your access to this project was removed. The unfinished upload was discarded.”you lost write accessaccess-lost (403)
“One upload is already in progress. This one continues when one finishes.”the deployment’s concurrent uploadsbusy (503, retried)
“This unfinished upload was removed after 7 days. Start it again.”the upload expiredexpired (410)

A volume the reader cannot use (not a full inline × crossline grid, an unsupported sample format) is refused with the reader’s own sentence, not a size sentence.

config.json limits takes the table’s keys; each is bounded by the schema and served by the deployment as it applies it:

"limits": {"upload_bytes": 33554432, "file_bytes": 4294967296, "upload_total_bytes": 17179869184,
"retained_bytes": 17179869184, "uploads_per_project": 10000, "grid_cells": 3000000}
KeyDefaultBounds
file_bytesupload_bytes1 KiB - 64 GiB
upload_total_bytes (every uploaded file and version, companion files, unfinished uploads)128 MiBup to 1 TiB
retained_bytes (all retained originals)512 MiBup to 1 TiB
uploads_per_project (uploaded files, each version counted, and unfinished uploads in one project), concurrent_uploads, finish_jobs100, 2, 1up to 100,000; 64; 16
session_seconds (unfinished upload kept)604800 (7 days)1 hour - 30 days
grid_cells (an uploaded grid)262144up to 4,194,304
max_traces, max_samples, max_inline_bytes (a volume)1,000,000; 20,000; 32 MiB

The packaged example allows 10,000 uploads per project rather than the default 100, so that a project with many wells has room for one file per well.

Disk. An upload needs free space of 2.5 times its size where the payloads are kept (the parts, the inline chunks and slack); a smaller remainder is refused before any byte is sent. /healthz answers 503 disk when the disk cannot hold 2.5 times file_bytes, and the packaged compose file’s health check reports the server unhealthy. Leave the qualification margin of 25 GB beside that.

Memory. Parts are verified and stored as they arrive and the check reads the file by ranges, so the gateway’s memory does not grow with the file. The packaged image is qualified with mem_limit: 1g (commented in compose.yaml). Measured on the packaged image, one operation at a time (gateway, above its idle of about 165 MiB):

OperationPayloads on the deployment’s diskPayloads in an S3 bucket
Upload of a 1.1 GiB volume (36 parts)154 MiB, 33 s392 MiB, 74 s
Upload of the 2.8 GB Q13a volume (83 parts)166 MiB, 67 s
One inline, crossline or time slice14-103 MiB14-70 MiB
A 64 MiB range of the original75-85 MiB98 MiB
A 2.8-million-cell grid, uploaded / read248 / 192 MiB

With an S3 store the check holds up to four 32 MiB parts while it reads the file, and each part read over HTTP is held twice for a moment; the figure was flat through the 1.1 GiB upload, because it is set by the part size, not the file. A slice’s memory follows the slice (bounded by max_inline_bytes). A grid is parsed whole, so its memory is bounded by grid_cells, not by file_bytes. Two uploads at once (concurrent_uploads) can add their parts’ memory together.

Backups carry the parts and chunks like any other payload; sweep_payloads keeps every part an original or an unfinished upload names.