fix(excel): fix segfault on xlsx with main content type declared via …#61
Conversation
…<Default> - Wrap the <Default> branch in iterate_files_by_contenttype_expat_callback_element_start (xlsxio_read.c) with #ifndef USE_MINIZIP so the minizip backend skips the zip directory traversal at compile time; previously the traversal ran inside an expat callback while [Content_Types].xml was open for streaming read, reentering the minizip single-state handle and crashing in unzGetCurrentFileInfo (upstream issue linuxdeepin#28, unfixed in xlsxio 0.2.36) - Add regression sample tests/file/test_xlsxio_default_crash.xlsx whose [Content_Types].xml declares the main content type via <Default> to cover the former crash path 修复(excel): 修复 main contenttype 经 <Default> 声明的 xlsx 解析段错误 - 在 iterate_files_by_contenttype_expat_callback_element_start 的 <Default> 分支用 #ifndef USE_MINIZIP 包裹,使 minizip 后端编译期跳过 zip 目录遍历;原实现该遍历在 expat 回调内执行,而此时 [Content_Types].xml 已打开流式读取,对同一 unzFile 重入导致 minizip 单状态机冲突,在 unzGetCurrentFileInfo 处段错误(上游 issue linuxdeepin#28,xlsxio 0.2.36 未修复) - 新增回归样本 tests/file/test_xlsxio_default_crash.xlsx,其 [Content_Types].xml 将 main contenttype 经 <Default> 声明,覆盖原崩溃路径 Log: 修复 xlsxio 在 minizip 后端下解析 [Content_Types].xml 中经 <Default> 声明的 main contenttype 时,因目录遍历重入已打开文件的流式读取状态而在 unzGetCurrentFileInfo 处段错误的问题,并补充回归样本 Task: https://pms.uniontech.com/task-view-391297.html brechtsanders/xlsxio#28
Reviewer's guide (collapsed on small PRs)Reviewer's GuideWrap the handler in iterate_files_by_contenttype_expat_callback_element_start with a USE_MINIZIP guard to avoid re-entering the minizip unzip state machine while [Content_Types].xml is being streamed, and add a regression XLSX sample exercising this former crash path. Sequence diagram for expat callback handling of Default with minizip versus libzipsequenceDiagram
participant Reader
participant Minizip
participant Libzip
participant Expat
participant IterateCallback
Reader->>Minizip: unzOpenCurrentFile
Reader->>Expat: expat_process_zip_file
Expat->>IterateCallback: iterate_files_by_contenttype_expat_callback_element_start
alt libzip_backend
IterateCallback->>Libzip: zip_get_name
end
alt minizip_backend_before_fix
IterateCallback->>Minizip: unzGoToFirstFile
IterateCallback->>Minizip: unzGetCurrentFileInfo
IterateCallback->>Minizip: unzGoToNextFile
end
alt minizip_backend_after_fix
Note over IterateCallback: [Default branch skipped when USE_MINIZIP]
end
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
|
Note
详情{
"3rdparty/libs/fileext/excel/xlsxio/xlsxio_read.c": [
{
"line": " if (XML_Char_icmp(reltype, X(\"http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet\")) == 0) {",
"line_number": 853,
"rule": "S35",
"reason": "Url link | 591ded0820"
},
{
"line": " } else if (XML_Char_icmp(reltype, X(\"http://schemas.openxmlformats.org/officeDocument/2006/relationships/sharedStrings\")) == 0) {",
"line_number": 861,
"rule": "S35",
"reason": "Url link | b67c93b15a"
},
{
"line": " } else if (XML_Char_icmp(reltype, X(\"http://schemas.openxmlformats.org/officeDocument/2006/relationships/styles\")) == 0) {",
"line_number": 866,
"rule": "S35",
"reason": "Url link | 7b6758f750"
}
]
} |
There was a problem hiding this comment.
Hey - I've left some high level feedback:
- The backend-specific logic is embedded directly in the XML callback with
#ifndef USE_MINIZIP; consider centralizing backend differences behind a small helper or abstraction so the parsing callback remains backend-agnostic and easier to reason about. - The new block comment around the
#ifndef USE_MINIZIPand the long#endiftrailing comment are quite verbose and language-mixed; consider shortening and standardizing them (e.g., English-only, with a brief explanation and an upstream issue reference) to keep the code easier to scan.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments
- The backend-specific logic is embedded directly in the XML callback with `#ifndef USE_MINIZIP`; consider centralizing backend differences behind a small helper or abstraction so the parsing callback remains backend-agnostic and easier to reason about.
- The new block comment around the `#ifndef USE_MINIZIP` and the long `#endif` trailing comment are quite verbose and language-mixed; consider shortening and standardizing them (e.g., English-only, with a brief explanation and an upstream issue reference) to keep the code easier to scan.Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.
deepin pr auto review★ 总体评分:100分■ 【总体评价】
■ 【详细分析】
■ 【改进建议代码示例】 --- a/3rdparty/libs/fileext/excel/xlsxio/xlsxio_read.c
+++ b/3rdparty/libs/fileext/excel/xlsxio/xlsxio_read.c
@@ -684,6 +684,13 @@ void iterate_files_by_contenttype_expat_callback_element_start (void* callbackda
}
} else if (XML_Char_icmp_ins(name, X("Default")) == 0) {
//by extension
+ // minizip 后端下:外层 expat_process_zip_file 已 unzOpenCurrentFile 打开
+ // [Content_Types].xml 并流式读取,此处对同一 unzFile 调用 unzGoToFirstFile /
+ // unzGetCurrentFileInfo / unzGoToNextFile 会破坏 minizip 单状态机导致段错误
+ // (上游 issue #28,xlsxio 0.2.36 仍未修复)。合法 xlsx 的 main contenttype 必
+ // 通过 <Override> 声明,<Default> 扩展名匹配对 xlsxio 无意义,故 minizip 后端
+ // 直接跳过本分支;libzip 后端基于索引的 zip_get_name 不受影响,逻辑保留。
+#ifndef USE_MINIZIP
const XML_Char* contenttype;
const XML_Char* extension;
if ((contenttype = get_expat_attr_by_name(atts, X("ContentType"))) != NULL && XML_Char_icmp(contenttype, data->contenttype) == 0) {
@@ -731,6 +738,7 @@ unzGetGlobalInfo(data->zip, &zipglobalinfo);
#endif
}
}
+#endif /* !USE_MINIZIP: 跳过 <Default> 分支,避免 minizip 状态冲突崩溃 */
}
} |
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: max-lvs, pppanghu77 The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:Approvers can indicate their approval by writing |
|
/forcemerge |
|
This pr force merged! (status: unstable) |
…
修复(excel): 修复 main contenttype 经 声明的 xlsx 解析段错误
Log: 修复 xlsxio 在 minizip 后端下解析 [Content_Types].xml 中经 声明的 main contenttype 时,因目录遍历重入已打开文件的流式读取状态而在 unzGetCurrentFileInfo 处段错误的问题,并补充回归样本
Task: https://pms.uniontech.com/task-view-391297.html brechtsanders/xlsxio#28
Summary by Sourcery
Prevent crashes when parsing XLSX content types with main type declared via using the minizip backend and add regression coverage.
Bug Fixes:
Tests: