java - lucene中增量更新的问题-6ren

java - lucene中增量更新的问题

转载作者：塔克拉玛干更新时间：2023-11-02 20:18:01

我正在创建一个程序，可以索引不同文件夹中的许多文本文件。所以这意味着每个包含文本文件的文件夹都会被索引，并且它的索引存储在另一个文件夹中。所以这个另一个文件夹就像我电脑中所有文件的通用索引。我正在使用 lucene 来实现这一点，因为 lucene 完全支持增量更新。这是我将其用于索引的源代码。

public class SimpleFileIndexer {


public static void main(String[] args) throws Exception   {

    int i=0;
    while(i<2) {
    File indexDir = new File("C:/Users/Raden/Documents/myindex");
    File dataDir = new File("C:/Users/Raden/Documents/indexthis");
    String suffix = "txt";

    SimpleFileIndexer indexer = new SimpleFileIndexer();

    int numIndex = indexer.index(indexDir, dataDir, suffix);

    System.out.println("Total files indexed " + numIndex);
    i++;
    Thread.sleep(1000);

    }
}


private int index(File indexDir, File dataDir, String suffix) throws Exception {
    RAMDirectory ramDir = new RAMDirectory();          // 1
    @SuppressWarnings("deprecation")
    IndexWriter indexWriter = new IndexWriter(
            ramDir,                                    // 2
            new StandardAnalyzer(Version.LUCENE_CURRENT),
            true,
            IndexWriter.MaxFieldLength.UNLIMITED);
    indexWriter.setUseCompoundFile(false);
    indexDirectory(indexWriter, dataDir, suffix);
    int numIndexed = indexWriter.maxDoc();
    indexWriter.optimize();
    indexWriter.close();

    Directory.copy(ramDir, FSDirectory.open(indexDir), false); // 3

    return numIndexed;
}


private void indexDirectory(IndexWriter indexWriter, File dataDir, String suffix)  throws IOException {
    File[] files = dataDir.listFiles();
    for (int i = 0; i < files.length; i++) {
        File f = files[i];
        if (f.isDirectory()) {
            indexDirectory(indexWriter, f, suffix);
        }
        else {
            indexFileWithIndexWriter(indexWriter, f, suffix);
        }
    }
}

private void indexFileWithIndexWriter(IndexWriter indexWriter, File f, String suffix) throws IOException {
    if (f.isHidden() || f.isDirectory() || !f.canRead() || !f.exists()) {
        return;
    }
    if (suffix!=null && !f.getName().endsWith(suffix)) {
        return;
    }
    System.out.println("Indexing file " + f.getCanonicalPath());

    Document doc = new Document();
    doc.add(new Field("contents", new FileReader(f)));      
doc.add(new Field("filename", f.getCanonicalPath(), Field.Store.YES, Field.Index.ANALYZED));
    indexWriter.addDocument(doc);
} }

这是我用来搜索 lucene 创建的索引的源代码

public class SimpleSearcher {

public static void main(String[] args) throws Exception {

    File indexDir = new File("C:/Users/Raden/Documents/myindex");
    String query = "revolution";
    int hits = 100;

    SimpleSearcher searcher = new SimpleSearcher();
    searcher.searchIndex(indexDir, query, hits);

}

private void searchIndex(File indexDir, String queryStr, int maxHits) throws Exception {

    Directory directory = FSDirectory.open(indexDir);

    IndexSearcher searcher = new IndexSearcher(directory);
    @SuppressWarnings("deprecation")
    QueryParser parser = new QueryParser(Version.LUCENE_30, "contents", new StandardAnalyzer(Version.LUCENE_CURRENT));
    Query query = parser.parse(queryStr);

    TopDocs topDocs = searcher.search(query, maxHits);

    ScoreDoc[] hits = topDocs.scoreDocs;
    for (int i = 0; i < hits.length; i++) {
        int docId = hits[i].doc;
        Document d = searcher.doc(docId);
        System.out.println(d.get("filename"));
    }

    System.out.println("Found " + hits.length);

}

}

我现在遇到的问题是我在上面创建的索引程序似乎无法进行任何增量更新。我的意思是我可以搜索一个文本文件，但只能搜索我已经索引到的最后一个文件夹中存在的文件，而我已经索引过的另一个先前的文件夹似乎在搜索结果中丢失并且没有显示.你能告诉我我的代码出了什么问题吗？我只是希望能够在我的源代码中具有增量更新功能。所以从本质上讲，我的程序似乎是用新索引覆盖现有索引而不是合并它。

谢谢

最佳答案

Directory.copy() 覆盖目标目录，您需要使用 IndexWriter.addIndexes() 将新目录索引合并到主目录中。

您也可以重新打开主索引并直接向其中添加文档。 RAMDirectory 不一定比适当调整缓冲区和合并因子设置更有效(参见 IndexWriter 文档)。

更新:代替Directory.copy()，您需要打开ramDir 进行读取，打开indexDir 进行写入和调用。在 indexDir 编写器上添加索引 并将其传递给 ramDir 读取器。或者，您可以使用 .addIndexesNoOptimize 并直接将其传递给 ramDir(无需打开阅读器)并在关闭前优化索引。

但实际上，跳过 RAMDir 并首先在 indexDir 上打开一个 writer 可能更容易。也将使更新更改的文件变得更加容易。

示例

private int index(File indexDir, File dataDir, String suffix) throws Exception {
    RAMDirectory ramDir = new RAMDirectory();
    IndexWriter indexWriter = new IndexWriter(ramDir,
        new StandardAnalyzer(Version.LUCENE_CURRENT), true,  
        IndexWriter.MaxFieldLength.UNLIMITED);
    indexWriter.setUseCompoundFile(false);
    indexDirectory(indexWriter, dataDir, suffix);
    int numIndexed = indexWriter.maxDoc();
    indexWriter.optimize();
    indexWriter.close();


    IndexWriter index = new IndexWriter(FSDirectory.open(indexDir),
        new StandardAnalyzer(Version.LUCENE_CURRENT), true,  
        IndexWriter.MaxFieldLength.UNLIMITED);
    index.addIndexesNoOptimize(ramDir);
    index.optimize();
    index.close();

    return numIndexed;
}

但是，这样也很好:

private int index(File indexDir, File dataDir, String suffix) throws Exception {

    IndexWriter index = new IndexWriter(FSDirectory.open(indexDir),
        new StandardAnalyzer(Version.LUCENE_CURRENT), true,  
        IndexWriter.MaxFieldLength.UNLIMITED);

    // tweak the settings for your hardware
    index.setUseCompoundFile(false);
    index.setRAMBufferSizeMB(256);
    index.setMergeFactor(30);

    indexDirectory(index, dataDir, suffix);

    index.optimize();
    int numIndexed = index.maxDoc();
    index.close();

    // you'll need to update indexDirectory() to keep track of indexed files
    return numIndexed;
}

关于java - lucene中增量更新的问题，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/4321221/

文章推荐： java - 如何下载具有Java依赖项的网页？

文章推荐： objective-c - ios:将指南针添加到 ARView

文章推荐： iOS - 初始加载大量数据

javascript - Mongoose 更新/更新？
我查看了网站上的一些问题，但还没有完全弄清楚我做错了什么。我有一些这样的代码: var mongoose = require('mongoose'), db = mongoose.connect('m
javascript - 更新、退出、更新、进入带有转换的模式
基本上，根据 this bl.ocks，我试图在开始新序列之前让所有 block 都变为 0。我认为我需要的是以下顺序: 更新为0 退出到0 更新随机数输入新号码我尝试通过添加以下代码块来遵循上述
java - 强制在线程内进行 GUI 更新 - JSlider 更新
我试图通过使用随机数在循环中设置 JSlider 位置来模拟“赛马”的投注结果。我的问题是，当然，我无法在线程执行时更新 GUI，因此我的 JSlider 似乎没有在竞赛，它们从头到尾都在运行。我尝试
php - PDO 更新帮助执行 pdo 更新
该功能非常简单: 变量:$table是正在更新的表$fields 是表中的字段，$values 从帖子生成并放入 $values 数组中而$where是表的索引字段的id值$indxfldnm 是索引
java - 数据库多线程插入(更新)和单线程顺序插入(更新)的性能比较？
让我们想象一个环境:有一个数据库客户端和一个数据库服务器。数据库客户端可以是 Java 程序或其他程序等；数据库服务器可以是mysql、oracle等。需求是在数据库服务器上的一个表中插入大量记录。
php - 更新、插入和删除时的 MySQL 更新 ID
在我当前的应用程序中，我正在制作一个菜单结构，它可以递归地创建自己的子菜单。然而，由于这个原因，我发现很难也允许某种重新排序方法。大多数应用程序可能只是通过“排序”列进行排序，但是在这种情况下，尽管这
ios - 更新/过期后供应配置文件 key 将更改 - 更新
Provisioning Profile 有 key ， key 链依赖于它。我想知道 key 什么时候会改变。 Key will change after renew Provisioning Pr
javascript - 是否应该发布 MongoDB 插入/更新/更新/删除？
截至目前，我在\server\publications.js 中有我的 MongoDB“选择”，例如: Meteor.publish("jobLocations", function () { r
ios - Swift:更新 UI - 主线程上的整个功能或只是 UI 更新？
我读到 UI 应该始终在主线程上更新。但是，当谈到实现这些更新的首选方法时，我有点困惑。我有各种函数可以执行一些条件检查，然后使用结果来确定如何更新 UI。我的问题是整个函数应该在主线程上运行吗？应
docker - yum 更新/apk 更新/apt-get 更新在代理后面不起作用
我在代理后面，我无法构建 Docker 镜像。我试过 FROM ubuntu , FROM centos和 FROM alpine ，但是 apt-get update/yum update/apk
java - 更新-更新 java truststore 中的自签名 CA 证书
我构建了一个 Java 应用程序，它向外部授权客户端公开网络服务。 Web 服务使用带有证书身份验证的 WS-security。基本上我们充当自定义证书颁发机构 - 我们在我们的服务器上维护一个 ja
asp.net - 更新 dll 时使用 app_offline.htm 使应用程序脱机更新 dll 时失败
因此，我有时会在上传新版本时使用 app_offline.htm 使应用程序离线。但是，当我上传较大的 dll 时，我收到黄色错误屏幕，指出无法加载 dll。这似乎与我对 app_offline.
visual-studio-cordova - 更新 Node 和 NPM VS Cordova 更新 5
我刚刚下载了 VS Apache Cordova Tools Update 5，但遇到了 Node 和 NPM 的问题。我使用默认的空白 cordova 项目进行测试。版本如果我在 VS 项目中对
angularjs - 避免 ng-view 在 $location.search 更新 GET 参数时获取 "wiped"(更新)
所以我有一个使用传单库实例化的 map 对象。 map 实例在单独的模板中创建并以这种方式路由:- var app = angular.module('myApp', ['ui', 'ngResour
java - Java 6 更新 19,20 中的绘图性能与 Java 6 更新 3 相比？
我使用较早的 Java 6 u 3 获得的帧速率是新版本的两倍。很奇怪。谁能解释一下？在 Core 2 Duo 1.83ghz 上，集成视频(仅使用一个内核)- 1500(较旧的 java)与 70
javascript - angular ng-click inside ng-repeat 更新 $scope 然后使用 $apply 更新 dom
我正在使用 angular 1.2 ng-repeat 创建的 div 也包含 ng-click 点击时 ng-click 更新 $scope $scope 中的变化反射(reflect)在使用 $a
android - public final void moveCamera(CameraUpdate 更新)和 public final void animateCamera(CameraUpdate 更新)之间的区别？
这些方法有什么区别 public final void moveCamera(CameraUpdate更新)和public final void animateCamera (CameraUpdate
列表树(更新)
我尝试了另一篇文章中某人评论中关于如何将树更改为列表的建议。但是，我在某处(或某物)有未声明的变量，所以我列表中的值是 [_G667, _G673, _G679]，而不是 [5, 2, 6]，这是正确
Java数据库大数据量查询/更新
实现以下场景的最佳方法是什么？我需要从java应用程序调用/查询包含数百万条记录的数据库表。然后，对于表中的每条记录，我的应用程序应该调用第三方 API 并获取状态字段作为响应。然后我的应用程序应该
Java重绘()/更新()
只是在编写一些与 java 图形相关的代码，这是我今天的讲座中的非常简单的示例。不管怎样，互联网似乎说更新不会被系统触发器调用，例如调整框架大小等。在这个例子中，更新是由这样的触发器调用的(因此当我只

塔克拉玛干

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

java - lucene中增量更新的问题