You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: apps/sim/lib/sim-search/live/README.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -32,7 +32,7 @@ For an explicit source user list, source-side verification tries those users' de
32
32
33
33
The source credential must be able to read the file. Selected folders, accessible subfolders, shared-drive IDs, and file types are checked using source-side metadata. Folder queries narrow candidate retrieval; metadata verification remains authoritative. A member's personal file outside the configured source scope is excluded even if their OAuth token can read it. Folder expansion and ancestry traversal are bounded and may report partial coverage.
34
34
35
-
Text-bearing PDF and DOCX files are downloaded with the member credential after checking `capabilities.canDownload`, then parsed with the shared document parsers. The download is capped at 4 MiB of decoded bytes; PDF extraction is complete or unavailable, with a 100-page and 1 MiB text limit. DOCX archives are checked before parsing (10 MiB total expanded, 4 MiB per entry), with a 1 MiB extracted-text limit. MIME/signature mismatches, password protection, malformed files, and documents with no extractable text fail explicitly; this path does not perform OCR. Existing source links and bounded comments/replies remain in successful reads. Other unsupported binary formats return labeled metadata only.
35
+
Text-bearing PDF and DOCX files are downloaded with the member credential after checking `capabilities.canDownload`, then parsed with the shared document parsers. The download is capped at 4 MiB of decoded bytes; PDF extraction is complete or unavailable, with a 100-page and 1 MiB text limit. DOCX archives are checked before parsing (10 MiB total expanded, 4 MiB per entry), with a 1 MiB extracted-text limit. Complete DOCX reads additionally bound the expanded conversion graph to 50,000 nodes and 2 MiB of model strings before HTML rendering, use built-in styles without embedding image bytes, and reject conversion failures instead of retrying a lossy fallback. MIME/signature mismatches, password protection, malformed files, and documents with no extractable text fail explicitly; this path does not perform OCR. Existing source links and bounded comments/replies remain in successful reads. Other unsupported binary formats return labeled metadata only.
0 commit comments