🖼️ feat: Return Sandbox Images From read_file as Viewable Artifacts (#14277)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run

* 🖼️ feat: Return Sandbox Images From `read_file` as Viewable Artifacts

The code-execution sandbox `read_file` path refused every image
extension because it reads files via `cat` over codeapi's JSON `/exec`
transport, which lossily corrupts non-UTF-8 bytes. The skill-file read
path already surfaced images as artifacts; this brings the sandbox path
to parity so an agent can actually see a chart/screenshot it reads.

- `readSandboxImage` (process.js): a Python base64 reader over `/exec`
  with an in-sandbox size guard so oversize images never cross the wire;
  base64 is ASCII-safe where `cat` corrupts.
- `handleSandboxImageRead` (handlers.ts): byte-integrity check (guards
  against a truncated `/exec` stdout), MIME resolved purely from the
  magic-byte sniff (extension only routes; a mislabeled non-image falls
  back to the bash hint), and graceful degradation on every failure mode.
- Shared `buildImageArtifactResult` used by both read paths; the result's
  `artifact.content` image_url reaches the UI (tool-end callbacks save it
  as an attachment) and the LLM (SDK folds it into the model-visible
  message for Anthropic/OpenAI/Google).

*  test: Sync read_file code-only description assertions with image wording

* 🛡️ fix: Harden sandbox image reads (regular-file guard, completeness check)

Addresses Codex review on PR #14277:

- readSandboxImage now os.stat's the target and rejects non-regular files
  (FIFOs, sockets, /dev/* symlinks) via stat.S_ISREG, and bounds the read at
  limit+1 bytes — a device/FIFO can no longer stream unbounded into memory
  until the request times out.
- handleSandboxImageRead validates completeness (not just the magic header):
  PNG must end with the IEND trailer and WebP's RIFF size must match the byte
  length, so a truncated/interrupted image degrades to the bash hint instead
  of being sent as a corrupt image_url. JPEG/GIF stay header-level (they can
  carry trailing metadata; a strict end-marker would risk false rejections).

* 🩹 fix: Chunk sandbox image reads to fit the runner stdout cap

Inlining any real image failed with "is an image file (.png) and cannot
be read as text". Root cause: readSandboxImage base64-encodes the file to
STDOUT, but the runner caps stdout at SANDBOX_OUTPUT_MAX_SIZE (1024 bytes
by default) and SIGKILLs the job on overflow (status OL), truncating the
JSON mid-base64. The parse then threw and the handler degraded to the
binary hint. The in-sandbox MAX_BINARY_BYTES=5MB guard never fired because
the *transport*, not the file size, is the real ceiling: a 5MB image needs
~6.8MB of stdout. Reproduced against a live MicroVM — a 186KB matplotlib
PNG died with 'stdout length exceeded' at exactly the 65536-byte cap.

Read the file in windows instead: each /exec pulls  raw bytes at an
offset and base64s only that slice, so every response stays under the cap
regardless of how the runner is configured; the chunks are reassembled and
verified against the sandbox-reported total. Verified end-to-end on a real
MicroVM: 25KB and 186KB PNGs both round-trip byte-exact (sha256 match).

Also:
- Detect the truncation explicitly (status OL) and name the fixable cause
  (chunk size / SANDBOX_OUTPUT_MAX_SIZE) instead of "unexpected output".
- Parse the LAST stdout line so a shell banner can't break the read, and
  include a stdout snippet when it genuinely is unparseable.
- LIBRECHAT_CODE_IMAGE_CHUNK_BYTES (default 32KB) tunes the window.
- Tests drive the real reader against a mocked /exec transport rather than
  mocking readSandboxImage, which is why the existing suite stayed green
  through this bug.

* 🎯 fix: Cap sandbox inline images at 1MB, separate from skill-file reads

The sandbox and skill-file image paths shared MAX_BINARY_BYTES (5MB), but
their transports differ: skill files stream from storage, while sandbox
bytes come back base64 over /exec stdout under the runner's output cap, so
the reader windows the file and cost scales in round-trips (~160 at 5MB vs
~32 at 1MB). Nothing is gained by allowing more — vision providers
downsample to ~1.5-2k px regardless, so multi-MB originals buy no fidelity
while grinding through round-trips.

Give the sandbox path its own MAX_SANDBOX_INLINE_IMAGE_BYTES (1MB), used
for both the read cap and the over-limit message (which previously quoted
5MB while the reader enforced something else). Skill-file reads keep 5MB.

Verified against a live MicroVM: a 186KB PNG round-trips byte-exact, and a
1.4MB file returns tooLarge in a single round-trip with zero bytes
transferred, degrading to the existing bash_tool hint.
This commit is contained in:
Danny Avila 2026-07-16 07:27:33 -04:00 committed by GitHub
parent 035228360d
commit 20cd00c492
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
7 changed files with 882 additions and 56 deletions

View file

@ -5,6 +5,7 @@ const {
getSessionInfo,
checkIfActive,
readSandboxFile,
readSandboxImage,
writeSandboxFile,
} = require('~/server/services/Files/Code/process');
const {
@ -358,6 +359,12 @@ const skillToolDeps = {
* the agents-side `ToolNode` via `tc.codeSessionContext`.
*/
readSandboxFile,
/**
* Companion to `readSandboxFile` for the raster-image case: pulls the
* bytes base64-encoded (size-guarded in-sandbox) so `read_file` can
* return an image the model can see instead of refusing it as binary.
*/
readSandboxImage,
writeSandboxFile,
};

View file

@ -1045,6 +1045,235 @@ async function readSandboxFile({ file_path, session_id, files, runtime_session_h
}
}
/**
* Reads a small image file out of the code-execution sandbox as base64 so
* `read_file` can surface it to vision-capable models. `readSandboxFile`'s
* `cat` round-trips stdout through codeapi's JSON transport, which lossily
* replaces non-UTF-8 bytes and corrupts image data. Here a tiny Python
* reader stats the file, refuses (without transferring) anything over
* `maxBytes`, and otherwise base64-encodes the bytes IN the sandbox so the
* payload stays ASCII-safe across the JSON `/exec` transport. Session
* forwarding mirrors `readSandboxFile` so the read lands in the same
* sandbox session that holds the agent's prior-turn artifacts.
*
* @param {Object} params
* @param {string} params.file_path - Path inside the sandbox (e.g. `/mnt/data/chart.png`).
* @param {string} [params.session_id] - Sandbox session id from the seeded context.
* @param {Array<{id: string, name: string, session_id?: string}>} [params.files] - File refs to mount.
* @param {string} [params.runtime_session_hint] - Per-conversation stateful runtime-session hint.
* @param {number} [params.maxBytes] - In-sandbox size cap; larger files return `{ tooLarge, bytes }`.
* @param {ServerRequest} [params.req] - Current authenticated request, used to mint Code API auth.
* @returns {Promise<{base64: string, bytes: number} | {tooLarge: true, bytes: number} | null>}
* `null` when codeapi is unavailable; throws on transport / read errors.
*/
async function readSandboxImage({
file_path,
session_id,
files,
runtime_session_hint,
maxBytes,
req,
}) {
const baseURL = getCodeBaseURL();
if (!baseURL) {
return null;
}
const limit = typeof maxBytes === 'number' && maxBytes > 0 ? maxBytes : 5 * megabyte;
const chunkBytes = getImageChunkBytes();
const maxChunks = Math.ceil(limit / chunkBytes) + 1;
/** @type {Buffer[]} */
const parts = [];
let offset = 0;
let total = null;
for (let i = 0; i < maxChunks; i++) {
const payload = Buffer.from(
JSON.stringify({ file_path, limit, offset, chunk: chunkBytes }),
'utf8',
).toString('base64');
const code = [
"python3 - <<'PY'",
'import base64, json, os, stat',
`payload = ${JSON.stringify(payload)}`,
"data = json.loads(base64.b64decode(payload).decode('utf-8'))",
"p = data['file_path']",
"limit = data['limit']",
"offset = data['offset']",
"chunk = data['chunk']",
'try:',
' st = os.stat(p)',
'except OSError as e:',
' print(json.dumps({"error": str(e)}))',
' raise SystemExit(0)',
// Reject FIFOs, sockets, and device files (e.g. a symlink to /dev/zero):
// os.stat can report a small/zero size while an unbounded read blocks or
// streams forever until the request times out.
'if not stat.S_ISREG(st.st_mode):',
' print(json.dumps({"error": "not a regular file"}))',
' raise SystemExit(0)',
'if st.st_size > limit:',
' print(json.dumps({"too_large": True, "bytes": st.st_size}))',
' raise SystemExit(0)',
// Read only this window. The whole base64 payload cannot be emitted in
// one shot: the runner caps stdout at SANDBOX_OUTPUT_MAX_SIZE (1024
// bytes by default) and SIGKILLs the job on overflow, which truncates
// the JSON mid-string. Windowing keeps every response under that cap.
"with open(p, 'rb') as f:",
' f.seek(offset)',
' raw = f.read(chunk)',
'print(json.dumps({"total": st.st_size, "n": len(raw), "b64": base64.b64encode(raw).decode("ascii")}))',
'PY',
].join('\n');
const parsed = await execSandboxImageChunk({
baseURL,
code,
file_path,
session_id,
runtime_session_hint,
files,
req,
chunkBytes,
});
if (parsed.error) {
throw new Error(String(parsed.error));
}
if (parsed.too_large === true) {
return { tooLarge: true, bytes: Number(parsed.bytes) || 0 };
}
if (typeof parsed.b64 !== 'string' || typeof parsed.n !== 'number') {
return null;
}
if (total == null) {
total = Number(parsed.total) || 0;
if (total > limit) {
return { tooLarge: true, bytes: total };
}
} else if (Number(parsed.total) !== total) {
/* The file changed underneath us; a spliced-together buffer would be
* a mix of two versions rather than any real image. */
throw new Error(`"${file_path}" changed while being read from the sandbox`);
}
parts.push(Buffer.from(parsed.b64, 'base64'));
offset += parsed.n;
if (parsed.n === 0 || offset >= total) {
break;
}
}
const buffer = Buffer.concat(parts);
if (total == null) {
return null;
}
if (buffer.length !== total) {
/* Ran out of chunk budget (or short reads); returning a partial image
* would render as a corrupt file, so surface it as unreadable-inline. */
return { tooLarge: true, bytes: total };
}
return { base64: buffer.toString('base64'), bytes: buffer.length };
}
/**
* Raw bytes pulled per `/exec` round-trip when inlining a sandbox image.
* Each chunk is base64-encoded (~1.33x) into the response's stdout, which
* the runner truncates + SIGKILLs past `SANDBOX_OUTPUT_MAX_SIZE`. The
* default leaves headroom for a runner configured at 64KB; deployments
* with a smaller cap must lower this, and a larger cap can raise it to cut
* round-trips.
* @returns {number}
*/
function getImageChunkBytes() {
const parsed = Number(process.env.LIBRECHAT_CODE_IMAGE_CHUNK_BYTES);
return Number.isFinite(parsed) && parsed > 0 ? Math.floor(parsed) : 32 * 1024;
}
/**
* Runs one image-chunk read over `/exec` and parses its JSON line.
* @returns {Promise<Record<string, unknown>>}
*/
async function execSandboxImageChunk({
baseURL,
code,
file_path,
session_id,
runtime_session_hint,
files,
req,
chunkBytes,
}) {
/** @type {Record<string, unknown>} */
const postData = { lang: 'bash', code };
if (session_id) {
postData.session_id = session_id;
}
if (runtime_session_hint) {
postData.runtime_session_hint = runtime_session_hint;
}
if (files && files.length > 0) {
postData.files = files;
}
try {
const authHeaders = await getCodeApiAuthHeaders(req);
const response = await axios({
method: 'post',
url: `${baseURL}/exec`,
data: postData,
headers: {
'Content-Type': 'application/json',
'User-Agent': 'LibreChat/1.0',
...authHeaders,
},
httpAgent: codeServerHttpAgent,
httpsAgent: codeServerHttpsAgent,
timeout: 15000,
});
const result = response?.data ?? {};
/* The runner truncates stdout at SANDBOX_OUTPUT_MAX_SIZE and SIGKILLs the
* job (status `OL`). Detect that explicitly: the surviving stdout is a
* base64 string cut mid-flight, so parsing it yields a misleading
* "unexpected output" instead of naming the real, fixable cause. */
if (result.status === 'OL') {
throw new Error(
`Reading "${file_path}" exceeded the sandbox stdout limit (chunk ${chunkBytes} bytes). ` +
'Lower LIBRECHAT_CODE_IMAGE_CHUNK_BYTES or raise SANDBOX_OUTPUT_MAX_SIZE on the runner.',
);
}
if (result.stderr && (result.stdout == null || result.stdout === '')) {
throw new Error(String(result.stderr).trim());
}
if (result.stdout == null || String(result.stdout).trim() === '') {
return {};
}
/* Parse the LAST non-empty line: the reader's JSON is the final thing it
* prints, so anything a shell profile or library emitted ahead of it
* (banners, warnings) must not break the read. */
const lines = String(result.stdout)
.split('\n')
.map((line) => line.trim())
.filter(Boolean);
try {
return JSON.parse(lines[lines.length - 1]);
} catch {
throw new Error(
`Unexpected output while reading image bytes from the sandbox: ${String(result.stdout).slice(0, 120)}`,
);
}
} catch (error) {
logAxiosError({
message: `Error reading sandbox image "${file_path}"`,
error,
});
throw error;
}
}
/**
* Writes a UTF-8 text file into the code-execution sandbox by running a
* small Python writer through the sandbox `/exec` endpoint. The payload is
@ -1150,6 +1379,7 @@ module.exports = {
getSessionInfo,
processCodeOutput,
readSandboxFile,
readSandboxImage,
writeSandboxFile,
runPreviewFinalize,
};

View file

@ -165,6 +165,7 @@ const {
processCodeOutput,
getSessionInfo,
readSandboxFile,
readSandboxImage,
writeSandboxFile,
primeFiles,
} = require('./process');
@ -2069,4 +2070,113 @@ describe('Code Process', () => {
expect(result.toolContext).not.toContain('preview');
});
});
/**
* These drive the REAL reader against a mocked `/exec` transport (rather
* than mocking `readSandboxImage` itself), because the bug this covers
* lived entirely in the transport: base64 leaves the sandbox on stdout,
* which the runner truncates + SIGKILLs past `SANDBOX_OUTPUT_MAX_SIZE`.
*/
describe('readSandboxImage', () => {
const crypto = require('crypto');
const PNG_HEADER = Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]);
/** Reply as the sandbox would: serve `buffer` through the windowed reader. */
const serveFile = (buffer) =>
mockAxios.mockImplementation(async ({ data }) => {
const payload = JSON.parse(
Buffer.from(JSON.parse(/payload = ("[^"]+")/.exec(data.code)[1]), 'base64').toString(),
);
const slice = buffer.subarray(payload.offset, payload.offset + payload.chunk);
return {
data: {
stdout: JSON.stringify({
total: buffer.length,
n: slice.length,
b64: slice.toString('base64'),
}),
},
};
});
beforeEach(() => {
process.env.LIBRECHAT_CODE_BASEURL = 'http://code.test/v1';
delete process.env.LIBRECHAT_CODE_IMAGE_CHUNK_BYTES;
mockAxios.mockReset();
});
it('reassembles an image larger than one chunk, byte-for-byte', async () => {
/* 200KB of PNG-headed noise: > 6 chunks at the 32KB default, and the
* exact shape that used to blow the stdout cap and SIGKILL the job. */
const source = Buffer.concat([PNG_HEADER, crypto.randomBytes(200 * 1024)]);
serveFile(source);
const result = await readSandboxImage({ file_path: '/mnt/data/big.png' });
expect(mockAxios.mock.calls.length).toBeGreaterThan(1);
expect(result.bytes).toBe(source.length);
expect(Buffer.from(result.base64, 'base64').equals(source)).toBe(true);
});
it('reads a single-chunk image in one round-trip', async () => {
const source = Buffer.concat([PNG_HEADER, crypto.randomBytes(1024)]);
serveFile(source);
const result = await readSandboxImage({ file_path: '/mnt/data/small.png' });
expect(mockAxios).toHaveBeenCalledTimes(1);
expect(Buffer.from(result.base64, 'base64').equals(source)).toBe(true);
});
it('names the real cause when a chunk overflows the runner stdout cap', async () => {
/* The runner truncates stdout and SIGKILLs with status `OL`; the old
* reader parsed the truncated base64 and reported a misleading
* "unexpected output" instead of the fixable limit. */
mockAxios.mockResolvedValue({
data: { stdout: '{"total":999999,"n":32768,"b64":"iVBORw0KGg', status: 'OL', code: 137 },
});
await expect(readSandboxImage({ file_path: '/mnt/data/big.png' })).rejects.toThrow(
/exceeded the sandbox stdout limit/,
);
});
it('honors LIBRECHAT_CODE_IMAGE_CHUNK_BYTES', async () => {
process.env.LIBRECHAT_CODE_IMAGE_CHUNK_BYTES = '1024';
const source = Buffer.concat([PNG_HEADER, crypto.randomBytes(4 * 1024)]);
serveFile(source);
const result = await readSandboxImage({ file_path: '/mnt/data/x.png' });
expect(mockAxios.mock.calls.length).toBe(5);
expect(Buffer.from(result.base64, 'base64').equals(source)).toBe(true);
});
it('parses the reader JSON even when the shell emits a banner first', async () => {
const source = Buffer.concat([PNG_HEADER, crypto.randomBytes(64)]);
mockAxios.mockResolvedValue({
data: {
stdout: `motd banner\n${JSON.stringify({
total: source.length,
n: source.length,
b64: source.toString('base64'),
})}`,
},
});
const result = await readSandboxImage({ file_path: '/mnt/data/x.png' });
expect(Buffer.from(result.base64, 'base64').equals(source)).toBe(true);
});
it('refuses an oversize file in-sandbox without transferring bytes', async () => {
mockAxios.mockResolvedValue({
data: { stdout: JSON.stringify({ too_large: true, bytes: 9 * 1024 * 1024 }) },
});
const result = await readSandboxImage({ file_path: '/mnt/data/huge.png' });
expect(result).toEqual({ tooLarge: true, bytes: 9 * 1024 * 1024 });
expect(mockAxios).toHaveBeenCalledTimes(1);
});
});
});

View file

@ -2361,6 +2361,7 @@ describe('createToolExecuteHandler', () => {
skillAuthoringAvailable?: boolean;
req?: unknown;
readSandboxFile?: ToolExecuteOptions['readSandboxFile'];
readSandboxImage?: ToolExecuteOptions['readSandboxImage'];
getSkillByName?: ToolExecuteOptions['getSkillByName'];
getAuthorSkillByName?: ToolExecuteOptions['getAuthorSkillByName'];
}) {
@ -2380,6 +2381,7 @@ describe('createToolExecuteHandler', () => {
getSkillByName: params.getSkillByName,
getAuthorSkillByName: params.getAuthorSkillByName,
readSandboxFile: params.readSandboxFile,
readSandboxImage: params.readSandboxImage,
});
}
@ -2964,16 +2966,214 @@ describe('createToolExecuteHandler', () => {
});
describe('binary file guard', () => {
/* 1x1 transparent PNG; decoded bytes start with the PNG magic so the
* handler's `sniffImageMime` resolves `image/png` regardless of the
* path extension. `pngBytes` feeds the integrity check that guards
* against codeapi truncating a large `/exec` stdout. */
const PNG_B64 =
'iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg==';
const pngBytes = Buffer.from(PNG_B64, 'base64').length;
/* Minimal magic-byte headers for the other supported formats enough
* for `sniffImageMime` to resolve the MIME from the actual bytes. The
* `read_file` MIME is always sniffed, never taken from the extension. */
const b64 = (bytes: number[]) => Buffer.from(bytes).toString('base64');
const JPEG_B64 = b64([0xff, 0xd8, 0xff, 0xe0, 0x00, 0x10]);
const GIF_B64 = b64([0x47, 0x49, 0x46, 0x38, 0x39, 0x61, 0x01, 0x00]);
/* RIFF container with a size field (bytes 4-7 LE = 12) that matches the
* 20-byte total, so the completeness check accepts it as intact. */
const WEBP_B64 = b64([
0x52, 0x49, 0x46, 0x46, 0x0c, 0, 0, 0, 0x57, 0x45, 0x42, 0x50, 0, 0, 0, 0, 0, 0, 0, 0,
]);
/* PNG magic header with NO IEND trailer — a truncated/interrupted write. */
const TRUNCATED_PNG_B64 = b64([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a, 0x00, 0x01]);
/* Plausible text bytes with no image magic — a file mislabeled `.png`. */
const NOT_IMAGE_B64 = Buffer.from('plainly text, not an image at all', 'utf8').toString(
'base64',
);
const imageUrlOf = (result: { artifact?: unknown }): string => {
const artifact = result.artifact as { content?: Array<{ image_url?: { url?: string } }> };
return artifact?.content?.[0]?.image_url?.url ?? '';
};
/**
* Regression for the matplotlib-shape bug where `read_file` on
* `/mnt/data/simple_graph.png` shelled `cat` through codeapi and
* line-numbered the lossy-string-decoded PNG bytes back to the
* model. The guard short-circuits BEFORE the network call for any
* extension that can never round-trip through codeapi's JSON
* `/exec` transport, and falls back to a NUL-byte sniff after the
* read for unknown extensions.
* `read_file` on a sandbox image returns the bytes as an `image_url`
* artifact the model can see. The SDK folds `artifact.content` into
* the model-visible message and the host tool-end callback saves the
* same data URL as a viewable attachment. `readSandboxFile` (the text
* `cat` path) must NOT be used its JSON transport corrupts image
* bytes, which was the matplotlib-shape mojibake regression.
*/
it('rejects images by extension without ever calling readSandboxFile', async () => {
it('returns a sandbox image as an image_url artifact the model can see', async () => {
const readSandboxFile = jest.fn();
const readSandboxImage = jest.fn(async () => ({ base64: PNG_B64, bytes: pngBytes }));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxFile,
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_png',
name: Constants.READ_FILE,
args: { path: '/mnt/data/simple_graph.png' },
codeSessionContext: { session_id: 'sess-Z', files: [] },
} as unknown as ToolCallRequest,
]);
expect(readSandboxFile).not.toHaveBeenCalled();
expect(readSandboxImage).toHaveBeenCalledWith(
expect.objectContaining({
file_path: '/mnt/data/simple_graph.png',
session_id: 'sess-Z',
maxBytes: expect.any(Number),
}),
);
expect(result.status).toBe('success');
expect(result.content).toContain('Image:');
expect(result.content).toContain('image/png');
expect(result.artifact).toMatchObject({
content: [{ type: 'image_url', image_url: { url: `data:image/png;base64,${PNG_B64}` } }],
});
});
it.each([
['png', '.png', PNG_B64, 'image/png'],
['jpeg', '.jpg', JPEG_B64, 'image/jpeg'],
['jpeg (.jpeg)', '.jpeg', JPEG_B64, 'image/jpeg'],
['gif', '.gif', GIF_B64, 'image/gif'],
['webp', '.webp', WEBP_B64, 'image/webp'],
])(
'inlines a %s image with the MIME sniffed from its bytes',
async (_label, ext, base64, expectedMime) => {
const bytes = Buffer.from(base64, 'base64').length;
const readSandboxImage = jest.fn(async () => ({ base64, bytes }));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: `call_${ext}`,
name: Constants.READ_FILE,
args: { path: `/mnt/data/asset${ext}` },
},
]);
expect(result.status).toBe('success');
expect(result.content).toContain(expectedMime);
expect(imageUrlOf(result)).toBe(`data:${expectedMime};base64,${base64}`);
},
);
it('declares the sniffed MIME, not the extension, when they disagree (.png holding JPEG bytes)', async () => {
/* matplotlib/PIL commonly re-encode to a different format than the
* filename suggests. The declared type must match the bytes or the
* provider rejects the image. */
const bytes = Buffer.from(JPEG_B64, 'base64').length;
const readSandboxImage = jest.fn(async () => ({ base64: JPEG_B64, bytes }));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_mismatch',
name: Constants.READ_FILE,
args: { path: '/mnt/data/actually_jpeg.png' },
},
]);
expect(result.status).toBe('success');
expect(imageUrlOf(result)).toBe(`data:image/jpeg;base64,${JPEG_B64}`);
});
it('refuses a non-image mislabeled with an image extension (bytes sniff to nothing)', async () => {
/* A renamed .txt/.pdf routed here by its `.png` name: the bytes match
* no supported image header, so we must NOT ship them declared as an
* image (the provider would reject) return the bash hint instead. */
const bytes = Buffer.from(NOT_IMAGE_B64, 'base64').length;
const readSandboxImage = jest.fn(async () => ({ base64: NOT_IMAGE_B64, bytes }));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_fake_png',
name: Constants.READ_FILE,
args: { path: '/mnt/data/notes.png' },
},
]);
expect(result.status).toBe('error');
expect(result.artifact).toBeUndefined();
expect(result.errorMessage).toContain('image file');
expect(result.errorMessage).toContain('bash_tool');
});
it('refuses a truncated image (valid magic header, missing trailer)', async () => {
/* A PNG whose write was interrupted keeps the magic prefix but lacks
* the IEND trailer; shipping it would fail saveBase64Image / the next
* provider request, so it must degrade to the bash hint. */
const bytes = Buffer.from(TRUNCATED_PNG_B64, 'base64').length;
const readSandboxImage = jest.fn(async () => ({ base64: TRUNCATED_PNG_B64, bytes }));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_truncated_png',
name: Constants.READ_FILE,
args: { path: '/mnt/data/half_written.png' },
},
]);
expect(result.status).toBe('error');
expect(result.artifact).toBeUndefined();
expect(result.errorMessage).toContain('image file');
expect(result.errorMessage).toContain('bash_tool');
});
it('routes to the image reader case-insensitively (.PNG)', async () => {
const readSandboxFile = jest.fn();
const readSandboxImage = jest.fn(async () => ({ base64: PNG_B64, bytes: pngBytes }));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxFile,
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_uppercase',
name: Constants.READ_FILE,
args: { path: '/mnt/data/CHART.PNG' },
},
]);
expect(readSandboxFile).not.toHaveBeenCalled();
expect(readSandboxImage).toHaveBeenCalledWith(
expect.objectContaining({ file_path: '/mnt/data/CHART.PNG' }),
);
expect(result.status).toBe('success');
expect(result.artifact).toBeDefined();
});
it('degrades to a bash-pointing image hint when no sandbox image reader is wired', async () => {
const readSandboxFile = jest.fn();
const handler = makeReadFileHandler({
codeEnvAvailable: true,
@ -2994,16 +3194,91 @@ describe('createToolExecuteHandler', () => {
expect(result.status).toBe('error');
expect(result.errorMessage).toContain('image file');
expect(result.errorMessage).toContain('.png');
expect(result.errorMessage).toContain('already attached');
expect(result.errorMessage).toContain('bash_tool');
expect(result.errorMessage).not.toContain('already attached');
});
it('reports an over-limit image without transferring bytes', async () => {
const readSandboxImage = jest.fn(async () => ({
tooLarge: true as const,
bytes: 9_000_000,
}));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_big',
name: Constants.READ_FILE,
args: { path: '/mnt/data/huge.png' },
},
]);
expect(result.status).toBe('success');
expect(result.artifact).toBeUndefined();
expect(result.content).toContain('inline limit');
expect(result.content).toContain('bash_tool');
});
it('degrades to the image hint when decoded bytes are truncated (integrity guard)', async () => {
/* Simulate codeapi clipping a large `/exec` stdout: the reported
* size does not match the decoded base64 length, so the bytes are
* unsafe to forward and we fall back to the bash hint. */
const readSandboxImage = jest.fn(async () => ({ base64: PNG_B64, bytes: pngBytes + 100 }));
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_trunc',
name: Constants.READ_FILE,
args: { path: '/mnt/data/clipped.png' },
},
]);
expect(result.status).toBe('error');
expect(result.artifact).toBeUndefined();
expect(result.errorMessage).toContain('image file');
expect(result.errorMessage).toContain('bash_tool');
});
it('rejects non-image binary types with a bash-pointing message (not the image-attachment hint)', async () => {
it('degrades to the image hint when the image reader throws', async () => {
const readSandboxImage = jest.fn(async () => {
throw new Error('codeapi unreachable');
});
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_throw',
name: Constants.READ_FILE,
args: { path: '/mnt/data/broken.png' },
},
]);
expect(result.status).toBe('error');
expect(result.errorMessage).toContain('image file');
expect(result.errorMessage).toContain('bash_tool');
});
it('rejects non-image binary types with a bash-pointing message (not the image path)', async () => {
const readSandboxFile = jest.fn();
const readSandboxImage = jest.fn();
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxFile,
readSandboxImage,
});
const [result] = await invokeHandler(handler, [
@ -3015,6 +3290,7 @@ describe('createToolExecuteHandler', () => {
]);
expect(readSandboxFile).not.toHaveBeenCalled();
expect(readSandboxImage).not.toHaveBeenCalled();
expect(result.status).toBe('error');
expect(result.errorMessage).toContain('binary file');
expect(result.errorMessage).toContain('.zip');
@ -3022,27 +3298,6 @@ describe('createToolExecuteHandler', () => {
expect(result.errorMessage).toContain('bash_tool');
});
it('is case-insensitive on the extension match (PNG vs .png)', async () => {
const readSandboxFile = jest.fn();
const handler = makeReadFileHandler({
codeEnvAvailable: true,
accessibleSkillIds: skillsInScope(),
readSandboxFile,
});
const [result] = await invokeHandler(handler, [
{
id: 'call_uppercase',
name: Constants.READ_FILE,
args: { path: '/mnt/data/CHART.PNG' },
},
]);
expect(readSandboxFile).not.toHaveBeenCalled();
expect(result.status).toBe('error');
expect(result.errorMessage).toContain('image file');
});
it('rejects binary content (NUL bytes) post-fetch when the extension was unknown', async () => {
/* No extension no precheck shortcut read goes through, but the
* NUL-byte sniff catches it before line-numbering. Mojibake from

View file

@ -253,6 +253,26 @@ export interface ToolExecuteOptions {
runtime_session_hint?: string;
req?: ServerRequest;
}) => Promise<{ content: string } | null>;
/**
* Reads a small image file out of the code-execution sandbox as base64 so
* `read_file` can surface it to vision-capable models. The `readSandboxFile`
* `cat` path round-trips stdout through codeapi's JSON transport, which
* lossily replaces non-UTF-8 bytes and mangles image data this reader
* base64-encodes the bytes IN the sandbox (ASCII-safe over JSON) after an
* in-sandbox size guard so an oversize image never crosses the wire.
* Returns `null` when codeapi is unavailable; throws on transport / read
* errors so the handler can fall back to an instructive message.
*/
readSandboxImage?: (params: {
file_path: string;
session_id?: string;
files?: Array<{ id: string; name: string; session_id?: string; storage_session_id?: string }>;
/** @see readSandboxFile.runtime_session_hint */
runtime_session_hint?: string;
/** In-sandbox size cap; files larger than this return `tooLarge` without transferring bytes. */
maxBytes?: number;
req?: ServerRequest;
}) => Promise<{ base64: string; bytes: number } | { tooLarge: true; bytes: number } | null>;
/**
* Writes a UTF-8 text file into the code-execution sandbox via the
* sandbox `/exec` endpoint. Mirrors `readSandboxFile` session forwarding
@ -277,6 +297,19 @@ export interface ToolExecuteOptions {
const MAX_READABLE_BYTES = 262_144;
const MAX_BINARY_BYTES = 5 * 1024 * 1024;
/**
* Inline ceiling for images pulled out of the code-execution sandbox
* deliberately tighter than {@link MAX_BINARY_BYTES}, which governs the
* skill-file path. The two differ because their transports differ: skill
* files stream from storage, while sandbox bytes come back base64 over
* `/exec` stdout, which the runner caps (`SANDBOX_OUTPUT_MAX_SIZE`). The
* reader therefore windows the file, so cost scales in round-trips
* ~32 at this limit vs ~160 at 5MB. Nothing is lost by stopping here:
* vision providers downsample to ~1.5-2k px regardless, so multi-MB
* originals buy no fidelity, and anything larger degrades to the
* `bash_tool` hint below.
*/
const MAX_SANDBOX_INLINE_IMAGE_BYTES = 1024 * 1024;
const MAX_CACHE_BYTES = 512 * 1024;
const MAX_AUTHORING_BYTES = 10 * 1024 * 1024;
const MAX_TOOL_ERROR_MESSAGE_CHARS = 12_000;
@ -1191,17 +1224,131 @@ function lowercaseExtension(filePath: string): string {
* Builds the model-visible error returned when `read_file` is invoked on
* a binary path. Phrasing is tuned for the LLM: states the fact (file is
* binary, can't be read as text), points at the correct affordance for
* each common case (image already in the chat; bash for everything else),
* and includes the path verbatim so the model can copy-paste into its
* next call.
* each common case (image via bash bytes; bash for everything else), and
* includes the path verbatim so the model can copy-paste into its next
* call. Supported raster images take the inline-attachment path first (see
* `handleSandboxImageRead`); this image branch is only reached when that
* read is unavailable (codeapi off) or fails.
*/
function buildBinaryFileError(filePath: string, ext: string): string {
if (IMAGE_EXTENSIONS_FOR_HINT.has(ext)) {
return `"${filePath}" is an image file (${ext}) and cannot be read as text. The image is already attached to the conversation and visible to the user. To process it programmatically, use \`bash_tool\` (e.g. \`file ${filePath}\` for metadata, or \`python3 -c '...'\` to operate on the bytes).`;
return `"${filePath}" is an image file (${ext}) and cannot be read as text. To process it programmatically, use \`bash_tool\` (e.g. \`file ${filePath}\` for metadata, or \`python3 -c '...'\` to operate on the bytes).`;
}
return `"${filePath}" is a binary file (${ext}) and cannot be read as text by \`read_file\`. Use \`bash_tool\` to process it (e.g. \`file ${filePath}\` for metadata, or a runtime-appropriate command for the format).`;
}
/**
* Sandbox file extensions `read_file` attempts to inline as visual content.
* The extension only decides ROUTING (try the base64 image read vs the text
* / bash path); the emitted MIME comes from the magic-byte sniff so the
* declared type always matches the actual bytes. Scoped to the four raster
* formats the providers accept in tool results (`IMAGE_MIMES`); other image
* extensions (`.bmp`, `.tiff`, `.svg`, ...) stay on the text / bash path.
*/
const SANDBOX_IMAGE_EXTENSIONS = new Set<string>(['.png', '.jpg', '.jpeg', '.gif', '.webp']);
/**
* Magic-byte sniff for the raster formats we inline. Preferred over the
* extension so a mislabelled `.png` that is really a JPEG is declared with
* the MIME the provider will actually validate the bytes against. Returns
* `undefined` when the header matches none of the supported formats.
*/
function sniffImageMime(buffer: Buffer): string | undefined {
if (buffer.length < 4) return undefined;
if (
buffer.length >= 8 &&
buffer[0] === 0x89 &&
buffer[1] === 0x50 &&
buffer[2] === 0x4e &&
buffer[3] === 0x47
) {
return 'image/png';
}
if (buffer[0] === 0xff && buffer[1] === 0xd8 && buffer[2] === 0xff) {
return 'image/jpeg';
}
if (
buffer.length >= 6 &&
buffer[0] === 0x47 &&
buffer[1] === 0x49 &&
buffer[2] === 0x46 &&
buffer[3] === 0x38
) {
return 'image/gif';
}
if (
buffer.length >= 12 &&
buffer[0] === 0x52 &&
buffer[1] === 0x49 &&
buffer[2] === 0x46 &&
buffer[3] === 0x46 &&
buffer[8] === 0x57 &&
buffer[9] === 0x45 &&
buffer[10] === 0x42 &&
buffer[11] === 0x50
) {
return 'image/webp';
}
return undefined;
}
/**
* Cheap structural check that the image bytes are complete, not just that the
* header sniffed valid a truncated/interrupted write can keep a valid magic
* prefix while the body is missing, which would then fail `saveBase64Image`
* resizing or the next provider request instead of the intended bash-hint
* fallback. Only png (fixed 8-byte IEND trailer) and webp (self-describing
* RIFF size) have a false-positive-free end marker; jpeg/gif can legitimately
* carry trailing metadata, so those stay at header-level sniffing rather than
* risk rejecting a valid file.
*/
function isCompleteImage(buffer: Buffer, mime: string): boolean {
if (mime === 'image/png') {
if (buffer.length < 8) return false;
const iend = buffer.subarray(buffer.length - 8);
return (
iend[0] === 0x49 &&
iend[1] === 0x45 &&
iend[2] === 0x4e &&
iend[3] === 0x44 &&
iend[4] === 0xae &&
iend[5] === 0x42 &&
iend[6] === 0x60 &&
iend[7] === 0x82
);
}
if (mime === 'image/webp') {
if (buffer.length < 12) return false;
return buffer.readUInt32LE(4) === buffer.length - 8;
}
return true;
}
/**
* Builds the `read_file` success result for an image: a short text line the
* model reads plus the `image_url` block in `artifact.content`. The SDK
* folds `artifact.content` into what the model sees (Anthropic tool_result
* or a trailing Human message for OpenAI/Google), and the host tool-end
* callback saves the same data URL as a viewable attachment. Shared by the
* skill-file and sandbox read paths so both surface images identically.
*/
function buildImageArtifactResult(
toolCallId: string,
displayPath: string,
mimeType: string,
bytes: number,
base64: string,
): ToolExecuteResult {
return {
toolCallId,
status: 'success',
content: `Image: ${displayPath} (${bytes} bytes, ${mimeType})`,
artifact: {
content: [{ type: 'image_url', image_url: { url: `data:${mimeType};base64,${base64}` } }],
},
};
}
/**
* True when the first chunk of a string contains a NUL byte. Used as a
* post-fetch safety net for files whose extension didn't match the
@ -1218,6 +1365,81 @@ function looksBinary(content: string): boolean {
return false;
}
/**
* Reads a sandbox image as a viewable artifact so `read_file` can hand the
* bytes to vision-capable models instead of refusing them. Fetches the file
* base64-encoded from the sandbox (`readSandboxImage`), verifies the decoded
* length matches the size the sandbox reported (guards against codeapi
* truncating a large `/exec` stdout into a corrupt image), sniffs the real
* MIME, and returns the shared image-artifact result. Degrades to the
* text-oriented binary hint when the reader is unavailable, the image is
* over the inline cap, or the read fails never throws.
*/
async function handleSandboxImageRead(
tc: ToolCallRequest,
filePath: string,
ext: string,
options: ToolExecuteOptions,
req?: ServerRequest,
): Promise<ToolExecuteResult> {
const { readSandboxImage } = options;
const binaryHint = (): ToolExecuteResult => ({
toolCallId: tc.id,
status: 'error',
content: '',
errorMessage: buildBinaryFileError(filePath, ext),
});
if (!readSandboxImage) {
return binaryHint();
}
const ctx = tc.codeSessionContext as SandboxSessionContext | undefined;
let read: { base64: string; bytes: number } | { tooLarge: true; bytes: number } | null;
try {
read = await readSandboxImage({
file_path: filePath,
session_id: ctx?.session_id,
files: ctx?.files,
maxBytes: MAX_SANDBOX_INLINE_IMAGE_BYTES,
...(tc.runtimeSessionHint ? { runtime_session_hint: tc.runtimeSessionHint } : {}),
...(req ? { req } : {}),
});
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
logger.warn(`[handleReadFileCall] Sandbox image read failed for "${filePath}": ${message}`);
return binaryHint();
}
if (!read) {
return binaryHint();
}
if ('tooLarge' in read) {
return {
toolCallId: tc.id,
status: 'success',
content: `Image "${filePath}" is ${read.bytes} bytes, over the ${MAX_SANDBOX_INLINE_IMAGE_BYTES}-byte inline limit. Use \`bash_tool\` to process it (e.g. \`file ${filePath}\` for metadata).`,
};
}
const buffer = Buffer.from(read.base64, 'base64');
if (buffer.length !== read.bytes) {
logger.warn(
`[handleReadFileCall] Sandbox image byte mismatch for "${filePath}" (decoded ${buffer.length} != reported ${read.bytes})`,
);
return binaryHint();
}
// Resolve the MIME from the actual bytes, never the extension: a file
// routed here by its `.png`/`.jpg`/... name whose header matches none of
// the supported formats is a mislabeled non-image (a renamed .txt/.pdf).
// Refuse it (and any truncated/incomplete image) with the bash hint
// instead of shipping bytes the provider would reject as a corrupt image.
const mimeType = sniffImageMime(buffer);
if (!mimeType || !isCompleteImage(buffer, mimeType)) {
return binaryHint();
}
return buildImageArtifactResult(tc.id, filePath, mimeType, buffer.length, read.base64);
}
/**
* Routes a `read_file` call to the code-execution sandbox via the
* host-provided `readSandboxFile` callback. The sandbox session id and
@ -1228,13 +1450,15 @@ function looksBinary(content: string): boolean {
* or an instructive error pointing the model at `bash_tool` when the
* sandbox isn't reachable from this configuration.
*
* Two binary guards keep `cat`-on-a-PNG-style mojibake out of the LLM
* context: (1) an extension precheck that short-circuits known binary
* types BEFORE any network call, and (2) a NUL-byte content sniff after
* the read for unknown extensions. The codeapi `/exec` transport is JSON,
* which already lossily down-converts non-UTF-8 stdout to replacement
* characters the bytes are unrecoverable here, so the goal is to fail
* fast with an instructive message rather than ship garbage.
* Supported raster images (`.png/.jpg/.jpeg/.gif/.webp`) take a dedicated
* base64 read path (`handleSandboxImageRead`) so the model can actually see
* them. Two binary guards then keep `cat`-on-a-PNG-style mojibake out of the
* LLM context for everything else: (1) an extension precheck that short-
* circuits known binary types BEFORE any network call, and (2) a NUL-byte
* content sniff after the read for unknown extensions. The codeapi `/exec`
* transport is JSON, which lossily down-converts non-UTF-8 `cat` stdout to
* replacement characters text bytes are unrecoverable there, so the goal
* is to fail fast with an instructive message rather than ship garbage.
*/
async function handleSandboxFileFallback(
tc: ToolCallRequest,
@ -1243,6 +1467,9 @@ async function handleSandboxFileFallback(
req?: ServerRequest,
): Promise<ToolExecuteResult> {
const ext = lowercaseExtension(filePath);
if (SANDBOX_IMAGE_EXTENSIONS.has(ext)) {
return handleSandboxImageRead(tc, filePath, ext, options, req);
}
if (BINARY_EXTENSIONS_NEVER_READABLE.has(ext)) {
return {
toolCallId: tc.id,
@ -2994,17 +3221,13 @@ async function handleReadFileCall(
// Return images/PDFs as artifacts
if (IMAGE_MIMES.has(file.mimeType) && buffer.length <= MAX_BINARY_BYTES) {
const base64 = buffer.toString('base64');
return {
toolCallId: tc.id,
status: 'success',
content: `Image: ${args.path} (${buffer.length} bytes, ${file.mimeType})`,
artifact: {
content: [
{ type: 'image_url', image_url: { url: `data:${file.mimeType};base64,${base64}` } },
],
},
};
return buildImageArtifactResult(
tc.id,
args.path,
file.mimeType,
buffer.length,
buffer.toString('base64'),
);
}
// TODO: PDF artifact support requires a document content block path

View file

@ -239,7 +239,8 @@ describe('registerCodeExecutionTools', () => {
expect(readFile?.description).toContain('/mnt/data/');
expect(readFile?.description).toContain('Do not run ls/find');
expect(readFile?.description).toContain('/tmp is per-call scratch');
expect(readFile?.description).toContain('truncated around 256KB');
expect(readFile?.description).toContain('truncate around 256KB');
expect(readFile?.description).toContain('images (png, jpeg, gif, webp)');
expect(readFile?.description).toContain('true filesystem discovery');
expect(readFile?.description).not.toContain('{skillName}');
expect(readFile?.description).not.toContain('SKILL.md');

View file

@ -140,9 +140,9 @@ const READ_FILE_DEF: LCTool = Object.freeze({
responseFormat: ReadFileToolDefinition.responseFormat,
}) as LCTool;
const CODE_READ_FILE_DESCRIPTION = `Read a known text file from the code-execution sandbox. Returns line-numbered text; large files may be truncated around 256KB.
const CODE_READ_FILE_DESCRIPTION = `Read a known file from the code-execution sandbox. Text files return line-numbered content (large files truncate around 256KB); images (png, jpeg, gif, webp) return as visual content you can see.
Use for text, CSV, JSON, Markdown, logs, and small source files at paths returned by tool output, just written, or under /mnt/data/. Do not run ls/find just to rediscover known paths. Use bash_tool for binary files, large files, transforms, metadata, or true filesystem discovery. /tmp is per-call scratch and unavailable later.`;
Use for text, CSV, JSON, Markdown, logs, small source files, and images at paths returned by tool output, just written, or under /mnt/data/. Do not run ls/find just to rediscover known paths. Use bash_tool for other binary files, large files, transforms, metadata, or true filesystem discovery. /tmp is per-call scratch and unavailable later.`;
const CODE_READ_FILE_PARAMETERS: LCTool['parameters'] = Object.freeze({
type: 'object',