gpt4 book ai didi

java - Apache POI 从文件读取时抛出编码错误,但从 Stream 读取时不会抛出编码错误

转载 作者:行者123 更新时间:2023-12-02 10:49:25 24 4
gpt4 key购买 nike

通过 Java (7) 中的 Apache POI (3.17) 加载特定 Excel (XLSX) 文件时,我收到有关编码的异常 (org.apache.xerces.impl.io.MalformedByteSequenceException: Invalid byte 2 4 字节 UTF-8 序列)。这似乎是在读取 sharedStrings.xml 文件时发生的(注意此文件采用 UTF8 编码)。

但是,如果我通过 InputStream 而不是 File 加载文件,则文件会正确加载。在这两种情况下,我都不会(或者我可以)指定编码。我知道从 InputStream 加载并不是最佳选择,我渴望避免这种情况。

我写了一个小例子来突出我的问题,但不幸的是我无法共享有问题的文件:

import java.io.File;
import java.io.FileInputStream;
import java.io.InputStream;
import java.text.MessageFormat;
import org.apache.poi.ss.usermodel.Workbook;
import org.apache.poi.ss.usermodel.WorkbookFactory;

public class POIEncodingIssue
{
public static void main(final String[] args) throws Exception
{
final File file = new File("path\\to\\my\\file.xlsx"); //$NON-NLS-1$
Workbook workbook = null;

// This works
System.out.println("Trying Stream based approach..."); //$NON-NLS-1$
try (InputStream stream = new FileInputStream(file))
{
workbook = WorkbookFactory.create(stream);

System.out.println(MessageFormat.format("Value was \"{0}\"", workbook.getSheetAt(0).getRow(0).getCell(0))); //$NON-NLS-1$
}
catch (final Exception e)
{
e.printStackTrace();
}
finally
{
if (workbook != null)
{
workbook.close();
}
}

// This doesn't
System.out.println("Trying File based approach..."); //$NON-NLS-1$
try
{
workbook = WorkbookFactory.create(file);

System.out.println(MessageFormat.format("Value was \"{0}\"", workbook.getSheetAt(0).getRow(0).getCell(0))); //$NON-NLS-1$
}
catch (final Exception e)
{
e.printStackTrace();
}
finally
{
if (workbook != null)
{
workbook.close();
}
}
}
}

以及产生的异常:

org.apache.poi.POIXMLException: java.lang.reflect.InvocationTargetException
at org.apache.poi.POIXMLFactory.createDocumentPart(POIXMLFactory.java:63)
at org.apache.poi.POIXMLDocumentPart.read(POIXMLDocumentPart.java:580)
at org.apache.poi.POIXMLDocument.load(POIXMLDocument.java:165)
at org.apache.poi.xssf.usermodel.XSSFWorkbook.<init>(XSSFWorkbook.java:270)
at org.apache.poi.ss.usermodel.WorkbookFactory.create(WorkbookFactory.java:266)
at org.apache.poi.ss.usermodel.WorkbookFactory.create(WorkbookFactory.java:226)
at org.apache.poi.ss.usermodel.WorkbookFactory.create(WorkbookFactory.java:205)
at com.in2.excelreader_art.DAT3983Example.main(DAT3983Example.java:42)
Caused by: java.lang.reflect.InvocationTargetException
at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)
at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)
at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)
at java.lang.reflect.Constructor.newInstance(Constructor.java:526)
at org.apache.poi.xssf.usermodel.XSSFFactory.createDocumentPart(XSSFFactory.java:56)
at org.apache.poi.POIXMLFactory.createDocumentPart(POIXMLFactory.java:60)
... 7 more
Caused by: java.io.IOException: Unable to parse xml bean
at org.apache.poi.POIXMLTypeLoader.parse(POIXMLTypeLoader.java:166)
at org.openxmlformats.schemas.spreadsheetml.x2006.main.SstDocument$Factory.parse(Unknown Source)
at org.apache.poi.xssf.model.SharedStringsTable.readFrom(SharedStringsTable.java:119)
at org.apache.poi.xssf.model.SharedStringsTable.<init>(SharedStringsTable.java:107)
... 13 more
Caused by: org.xml.sax.SAXParseException; lineNumber: 2; columnNumber: 5443012; Invalid byte 2 of 4-byte UTF-8 sequence.
at org.apache.xerces.util.ErrorHandlerWrapper.createSAXParseException(Unknown Source)
at org.apache.xerces.util.ErrorHandlerWrapper.fatalError(Unknown Source)
at org.apache.xerces.impl.XMLErrorReporter.reportError(Unknown Source)
at org.apache.xerces.impl.XMLErrorReporter.reportError(Unknown Source)
at org.apache.xerces.impl.XMLDocumentFragmentScannerImpl$FragmentContentDispatcher.dispatch(Unknown Source)
at org.apache.xerces.impl.XMLDocumentFragmentScannerImpl.scanDocument(Unknown Source)
at org.apache.xerces.parsers.XML11Configuration.parse(Unknown Source)
at org.apache.xerces.parsers.XML11Configuration.parse(Unknown Source)
at org.apache.xerces.parsers.XMLParser.parse(Unknown Source)
at org.apache.xerces.parsers.DOMParser.parse(Unknown Source)
at org.apache.xerces.jaxp.DocumentBuilderImpl.parse(Unknown Source)
at javax.xml.parsers.DocumentBuilder.parse(DocumentBuilder.java:121)
at org.apache.poi.util.DocumentHelper.readDocument(DocumentHelper.java:140)
at org.apache.poi.POIXMLTypeLoader.parse(POIXMLTypeLoader.java:163)
... 16 more
Caused by: org.apache.xerces.impl.io.MalformedByteSequenceException: Invalid byte 2 of 4-byte UTF-8 sequence.
at org.apache.xerces.impl.io.UTF8Reader.invalidByte(Unknown Source)
at org.apache.xerces.impl.io.UTF8Reader.read(Unknown Source)
at org.apache.xerces.impl.XMLEntityScanner.load(Unknown Source)
at org.apache.xerces.impl.XMLEntityScanner.scanContent(Unknown Source)
at org.apache.xerces.impl.XMLDocumentFragmentScannerImpl.scanContent(Unknown Source)
... 26 more

最佳答案

将评论提升为答案

您似乎在旧版本的 Apache XML Beans 中遇到了错误。如果您至少升级到Apache XML Beans 3.0.1您应该会发现问题消失了。

理想情况下,您还应该至少升级到 Apache POI 4.0.0,这需要较新的 xmlbean,但这需要 Java 8+。 XML Beans 向后兼容,因此您可以坚持使用 POI 3.17 并仅升级 xmlbeans,不会出现任何问题(尽管显然没有 4 中的 POI 修复!)

关于java - Apache POI 从文件读取时抛出编码错误,但从 Stream 读取时不会抛出编码错误,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/52283752/

24 4 0