java - JVM 抛出 java.lang.OutOfMemoryError : heap space (File processing)-6ren

java - JVM 抛出 java.lang.OutOfMemoryError : heap space (File processing)

转载作者：行者123 更新时间：2023-12-01 19:58:09

26

4

我编写了一个文件重复处理器，它获取每个文件的 MD5 哈希值，将其添加到 HashMap 中，然后获取具有相同哈希值的所有文件并将其添加到名为 dupeList 的 HashMap 中。但是在运行大型目录进行扫描时，例如 C:\Program Files\，它将抛出以下错误

Exception in thread "main" java.lang.OutOfMemoryError: Java heap space
at java.nio.file.Files.read(Unknown Source)
at java.nio.file.Files.readAllBytes(Unknown Source)
at com.embah.FileDupe.Utils.FileUtils.getMD5Hash(FileUtils.java:14)
at com.embah.FileDupe.FileDupe.getDuplicateFiles(FileDupe.java:43)
at com.embah.FileDupe.FileDupe.getDuplicateFiles(FileDupe.java:68)
at ImgHandler.main(ImgHandler.java:14)

我确信这是因为它处理如此多的文件，但我不确定有更好的方法来处理它。我正在努力让它发挥作用，这样我就可以筛选所有 child 的婴儿照片并删除重复的照片，然后再将它们放在外部硬盘上进行长期存储。谢谢大家的帮助!

我的代码

public class FileUtils {
public static String getMD5Hash(String path){
    try {
        byte[] bytes = Files.readAllBytes(Paths.get(path)); //LINE STACK THROWS ERROR
        byte[] hash = MessageDigest.getInstance("MD5").digest(bytes);
        bytes = null;
        String hexHash = DatatypeConverter.printHexBinary(hash);
        hash = null;
        return hexHash;
    } catch(Exception e){
        System.out.println("Having problem with file: " + path);
        return null;
    }
}

public class FileDupe {
public static Map<String, List<String>> getDuplicateFiles(String dirs){
    Map<String, List<String>> allEntrys = new HashMap<>(); //<hash, file loc>
    Map<String, List<String>> dupeEntrys = new HashMap<>();
    File fileDir = new File(dirs);
    if(fileDir.isDirectory()){
        ArrayList<File> nestedFiles = getNestedFiles(fileDir.listFiles());
        File[] fileList = new File[nestedFiles.size()];
        fileList = nestedFiles.toArray(fileList);

        for(File file:fileList){
            String path = file.getAbsolutePath();
            String hash = "";
            if((hash = FileUtils.getMD5Hash(path)) == null)
                continue;
            if(!allEntrys.containsValue(path))
                put(allEntrys, hash, path);
        }
        fileList = null;
    }
    allEntrys.forEach((hash, locs) -> {
        if(locs.size() > 1){
            dupeEntrys.put(hash, locs);
        }
    });
    allEntrys = null;
    return dupeEntrys;
}

public static Map<String, List<String>> getDuplicateFiles(String... dirs){
    ArrayList<Map<String, List<String>>> maps = new ArrayList<Map<String, List<String>>>();
    Map<String, List<String>> dupeMap = new HashMap<>();
    for(String dir : dirs){ //Get all dupe files
        maps.add(getDuplicateFiles(dir));
    }
    for(Map<String, List<String>> map : maps){ //iterate thru each map, and add all items not in the dupemap to it
        dupeMap.putAll(map);
    }
    return dupeMap;
}

protected static ArrayList<File> getNestedFiles(File[] fileDir){
    ArrayList<File> files = new ArrayList<File>();
    return getNestedFiles(fileDir, files);
}

protected static ArrayList<File> getNestedFiles(File[] fileDir, ArrayList<File> allFiles){
    for(File file:fileDir){
        if(file.isDirectory()){
            getNestedFiles(file.listFiles(), allFiles);
        } else {
            allFiles.add(file);
        }
    }
    return allFiles;
}

protected static <KEY, VALUE> void put(Map<KEY, List<VALUE>> map, KEY key, VALUE value) {
    map.compute(key, (s, strings) -> strings == null ? new ArrayList<>() : strings).add(value);
}


public class ImgHandler {
private static Scanner s = new Scanner(System.in);

public static void main(String[] args){
    System.out.print("Please enter locations to scan for dupelicates\nSeperate Location via semi-colon(;)\nLocations: ");
    String[] locList = s.nextLine().split(";");
    Map<String, List<String>> dupes = FileDupe.getDuplicateFiles(locList);
    System.out.println(dupes.size() + " dupes detected!");
    dupes.forEach((hash, locs) -> {
        System.out.println("Hash: " + hash);
        locs.forEach((loc) -> System.out.println("\tLocation: " + loc));
    });
}

最佳答案

将整个文件读入字节数组不仅需要足够的堆空间，原则上还限制文件大小最大为 Integer.MAX_VALUE(实际限制对于 HotSpot JVM 来说甚至还小了几个字节)。

最好的解决方案是根本不将数据加载到堆内存中:

public static String getMD5Hash(String path) {
    MessageDigest md;
    try { md = MessageDigest.getInstance("MD5"); }
    catch(NoSuchAlgorithmException ex) {
        System.out.println("FileUtils.getMD5Hash(): "+ex);
        return null;// TODO better error handling
    }
    try(FileChannel fch = FileChannel.open(Paths.get(path), StandardOpenOption.READ)) {
        for(long pos = 0, rem = fch.size(), chunk; rem>pos; pos+=chunk) {
            chunk = Math.min(Integer.MAX_VALUE, rem-pos);
            md.update(fch.map(FileChannel.MapMode.READ_ONLY, pos, chunk));
        }
    } catch(IOException e){
        System.out.println("Having problem with file: " + path);
        return null;// TODO better error handling
    }
    return String.format("%032X", new BigInteger(1, md.digest()));
}

如果底层的 MessageDigest 实现是一个纯 Java 实现，它会将数据从直接缓冲区传输到堆，但这超出了您的责任(并且这将是一个合理的权衡)消耗的堆内存和性能)。

上述方法可以毫无问题地处理超过 2GiB 大小的文件。

关于java - JVM 抛出 java.lang.OutOfMemoryError : heap space (File processing)，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/48777834/

26

4

0

文章推荐： ios - UIWebView显示警报

文章推荐： java - 复制 HashMap 的线程安全方法

文章推荐： java - 如何使用 POSTMAN JSON 发布 "Date"？

Grails 启动错误 : groovy. lang.Closure.rehydrate(Ljava/lang/Object;Ljava/lang/Object;Ljava/lang/Object;)Lgroovy/lang/Closure;
在 Tomcat 6/Ubuntu 12.04 上启动 Grails 2.1.0 应用程序时出现以下错误。 Error 500 - Internal Server Error. groovy.lang
java.lang.RuntimeException : java. lang.ClassCastException : java. lang.Long 无法转换为 java.lang.String
在运行 Storm 拓扑时，我收到此错误。拓扑完美运行 5 分钟，没有任何错误，然后失败。我正在使用 Config.TOPOLOGY_TICK_TUPLE_FREQ_SECS as 300 sec i
java.lang.NoSuchMethodException : java. lang.String.substring(java.lang.Long，java.lang.Long)
我有一个 jsp 代码在其中一台机器上运行良好。但是当我复制到另一台机器时，我得到了这个 no such method found 异常。我是 Spring 的新手。有人可以解释我错过了什么吗？以下
java - java.util.Map 中的 put(java.lang.String,java.lang.Inte ger) 不能应用于 (java.lang.String,java.lang.String )
已关闭。此问题需要 debugging details 。目前不接受答案。编辑问题以包含 desired behavior, a specific problem or error, and the
java.lang.NoSuchMethodError - Ljava/lang/String;)Ljava/lang/String;
我的代码在下面给出了一个错误； Exception in thread "main" java.lang.NoSuchMethodError: com/myApp/Client.cypherCBC(L
java - Restful Web 服务 - java.lang.NoSuchMethodError concat(Ljava/lang/Iterable;Ljava/lang/Iterable;)Ljava/lang/Iterable
我正在尝试一个 Restful web 服务示例，所以当我要访问 url 时，我遇到了异常 java.lang.NoSuchMethodError: jersey.repackaged.com.goo
java.lang.NoSuchMethodError : org. jboss.logging.Logger.getMessageLogger(Ljava/lang/Class;Ljava/lang/String;)Ljava/lang/Object;
我正在将一个 Spring web 项目转换为一个 Maven 项目，但我收到了这个错误: java.lang.NoSuchMethodError: org.jboss.logging.Logger.
java.lang.ClassCastException : java. lang.Short 无法转换为 java.lang.Integer
在我的项目中，我有一个像这样的枚举: public enum MyEnum { FIRST(1), SECOND(2); private int value; private MyEnum(int v
java.lang.ClassCastException : java. lang.String 无法转换为 java.lang.Long
我创建了这个简单的示例，用于读取 Linux 正常运行时间: public String getMachineUptime() throws IOException { String[] di
java.lang.ClassCastException : [Ljava. lang.Object;无法转换为 [Ljava.lang.Comparable;
我正在使用 Eclipse，并且正在使用 Java。我的目标是使用 bogoSort 方法对 vector 进行排序在一个 vector (vectorExample)中适应我的 vector 类型，
java.lang.ClassCastException : java. lang.String 无法转换为 [Ljava.lang.Object
我正在运行以下查询。它显示一条错误消息。如何解决这个错误？ ListrouteList=null; List companyList = session.createS
java.lang.ClassCastException : [Ljava. lang.Long;无法转换为 java.lang.Long
我有以下模型类: @Entity @Table(name="user_content") @org.hibernate.annotations.NamedQueries({ @org.
java.lang.ClassCastException : [Ljava. lang.String;无法转换为 java.lang.String
我有那个错误。这是我的代码: GmailSettingsService service = new GmailSettingsService(APPLICATION_NAME, DOMAIN_NAME
java.lang.ClassCastException : java. lang.StackOverflowError 无法转换为 java.lang.Exception
实际上我在执行我的java程序时遇到了下面提到的错误 Exception in thread "pool-1-thread-1" java.lang.ClassCastException: jav
java.lang.ClassCastException : java. lang.Float 无法转换为 java.lang.String
java.lang.ClassCastException: java.lang.Float cannot be cast to java.lang.String 我在以下代码中遇到此异常: Strin
java.lang.ClassCastException : [Ljava. lang.Object;不兼容 [Ljava.lang.String;
我正在尝试从 linkedhashset 中检索随机元素。下面是我的代码，但它每次都给我异常。 private static void generateRandomUserId(Set userIds
java.lang.ClassCastException : java. lang.Double 无法转换为 java.lang.String
我已经完成了 Android 中的代码: List spinnerArray = new ArrayList(); for (int i = 0; i item = (LinkedTreeMap)
java.lang.ClasscastException : java. lang.Long 无法转换为 java.lang.Double
这个问题已经有答案了: Explanation of ClassCastException in Java (12 个回答) 已关闭 6 年前。我已经编写了 java 到 Json 的代码，同时从页
java.lang.ClassCastException : java. lang.Long 无法转换为 clojure.lang.IFn
这个问题在这里已经有了答案: ClassCastException java.lang.Long cannot be cast to clojure.lang.IFn (4 个答案) 关闭 6 年前
java.lang.ClassCastException : java. lang.Integer 无法转换为 java.lang.Double
我在运行时遇到问题来编译这段代码，这给我一个错误，java.lang.Integer 无法转换为 Java.lang.Double。如果有人帮助我更正此代码，我将非常高兴 double x; pu

首页

博学

6Ren·AI

商城

java - JVM 抛出 java.lang.OutOfMemoryError : heap space (File processing)