django - django-haystack自动完成返回的结果太宽-6ren

django - django-haystack自动完成返回的结果太宽

转载作者：行者123 更新时间：2023-12-03 00:33:38

我用字段title_auto创建了一个索引:

class GameIndex(indexes.SearchIndex, indexes.Indexable):
    text = indexes.CharField(document=True, model_attr='title')
    title = indexes.CharField(model_attr='title')
    title_auto = indexes.NgramField(model_attr='title')

flex 搜索设置如下所示:

ELASTICSEARCH_INDEX_SETTINGS = {
    'settings': {
        "analysis": {
            "analyzer": {
                "ngram_analyzer": {
                    "type": "custom",
                    "tokenizer": "lowercase",
                    "filter": ["haystack_ngram"],
                    "token_chars": ["letter", "digit"]
                },
                "edgengram_analyzer": {
                    "type": "custom",
                    "tokenizer": "lowercase",
                    "filter": ["haystack_edgengram"]
                }
            },
            "tokenizer": {
                "haystack_ngram_tokenizer": {
                    "type": "nGram",
                    "min_gram": 1,
                    "max_gram": 15,
                },
                "haystack_edgengram_tokenizer": {
                    "type": "edgeNGram",
                    "min_gram": 1,
                    "max_gram": 15,
                    "side": "front"
                }
            },
            "filter": {
                "haystack_ngram": {
                    "type": "nGram",
                    "min_gram": 1,
                    "max_gram": 15
                },
                "haystack_edgengram": {
                    "type": "edgeNGram",
                    "min_gram": 1,
                    "max_gram": 15
                }
            }
        }
    }
}

我尝试进行自动完成搜索，但是可以返回太多不相关的结果:
qs = SearchQuerySet().models(Game).autocomplete(title_auto=search_phrase)
要么
qs = SearchQuerySet().models(Game).filter(title_auto=search_phrase)
它们都产生相同的输出。

如果search_phrase为“monopoly”，则第一个结果的标题中包含“Monopoly”，但是，由于只有2个相关项，因此返回51。其他与“Monopoly”无关。

所以我的问题是-如何更改结果的相关性？

最佳答案

由于我还没有看到完整的映射，因此很难确定，但是我怀疑问题是分析器(其中之一)同时用于索引和搜索。因此，当您为文档建立索引时，会创建并索引许多ngram术语。如果您搜索并且对搜索文本也进行了相同的分析，则会生成许多搜索词。由于最小的ngram是单个字母，因此几乎所有查询都将匹配许多文档。

我们在博客文章http://blog.qbox.io/multi-field-partial-word-autocomplete-in-elasticsearch-using-ngrams上写了一篇关于将ngrams用于自动完成的文章，您可能会觉得有帮助。但是，我将给您提供一个更简单的示例来说明我的意思。我对干草堆不是很熟悉，所以我可能无法为您提供帮助，但是我可以在Elasticsearch中用ngrams解释问题。

首先，我将建立一个使用ngram分析器进行索引和搜索的索引:

PUT /test_index
{
   "settings": {
       "number_of_shards": 1,
      "analysis": {
         "filter": {
            "nGram_filter": {
               "type": "nGram",
               "min_gram": 1,
               "max_gram": 15,
               "token_chars": [
                  "letter",
                  "digit",
                  "punctuation",
                  "symbol"
               ]
            }
         },
         "analyzer": {
            "nGram_analyzer": {
               "type": "custom",
               "tokenizer": "whitespace",
               "filter": [
                  "lowercase",
                  "asciifolding",
                  "nGram_filter"
               ]
            }
         }
      }
   },
   "mappings": {
        "doc": {
            "properties": {
                "title": {
                    "type": "string", 
                    "analyzer": "nGram_analyzer"
                }
            }
        }
   }
}

并添加一些文档:

PUT /test_index/_bulk
{"index":{"_index":"test_index","_type":"doc","_id":1}}
{"title":"monopoly"}
{"index":{"_index":"test_index","_type":"doc","_id":2}}
{"title":"oligopoly"}
{"index":{"_index":"test_index","_type":"doc","_id":3}}
{"title":"plutocracy"}
{"index":{"_index":"test_index","_type":"doc","_id":4}}
{"title":"theocracy"}
{"index":{"_index":"test_index","_type":"doc","_id":5}}
{"title":"democracy"}

并运行一个简单的 match搜索 "poly":

POST /test_index/_search
{
    "query": {
        "match": {
           "title": "poly"
        }
    }
}

它返回所有五个文档:

{
   "took": 3,
   "timed_out": false,
   "_shards": {
      "total": 1,
      "successful": 1,
      "failed": 0
   },
   "hits": {
      "total": 5,
      "max_score": 4.729521,
      "hits": [
         {
            "_index": "test_index",
            "_type": "doc",
            "_id": "2",
            "_score": 4.729521,
            "_source": {
               "title": "oligopoly"
            }
         },
         {
            "_index": "test_index",
            "_type": "doc",
            "_id": "1",
            "_score": 4.3608603,
            "_source": {
               "title": "monopoly"
            }
         },
         {
            "_index": "test_index",
            "_type": "doc",
            "_id": "3",
            "_score": 1.0197333,
            "_source": {
               "title": "plutocracy"
            }
         },
         {
            "_index": "test_index",
            "_type": "doc",
            "_id": "4",
            "_score": 0.31496215,
            "_source": {
               "title": "theocracy"
            }
         },
         {
            "_index": "test_index",
            "_type": "doc",
            "_id": "5",
            "_score": 0.31496215,
            "_source": {
               "title": "democracy"
            }
         }
      ]
   }
}

这是因为搜索项 "poly"被标记为术语 "p"， "o"， "l"和 "y"，由于每个文档中的 "title"字段被标记为单字母术语，因此它们与每个文档匹配。

如果我们改用此映射重建索引(相同的分析器和文档):

"mappings": {
  "doc": {
     "properties": {
        "title": {
           "type": "string",
           "index_analyzer": "nGram_analyzer",
           "search_analyzer": "standard"
        }
     }
  }
}

该查询将返回我们期望的结果:

POST /test_index/_search
{
    "query": {
        "match": {
           "title": "poly"
        }
    }
}
...
{
   "took": 1,
   "timed_out": false,
   "_shards": {
      "total": 1,
      "successful": 1,
      "failed": 0
   },
   "hits": {
      "total": 2,
      "max_score": 1.5108256,
      "hits": [
         {
            "_index": "test_index",
            "_type": "doc",
            "_id": "1",
            "_score": 1.5108256,
            "_source": {
               "title": "monopoly"
            }
         },
         {
            "_index": "test_index",
            "_type": "doc",
            "_id": "2",
            "_score": 1.5108256,
            "_source": {
               "title": "oligopoly"
            }
         }
      ]
   }
}

边缘ngram的工作原理类似，除了仅使用单词开头的术语。

这是我用于此示例的代码:

http://sense.qbox.io/gist/b24cbc531b483650c085a42963a49d6a23fa5579

关于django - django-haystack自动完成返回的结果太宽，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/29008725/

文章推荐： elasticsearch - 跨集群的ElasticSearch快照

文章推荐： powershell - 根据Get-AdUser的结果设置AD用户的UPN

文章推荐： spring - Spring Boot ElasticSearch端口

dart - 对于 "pubsub.stream.listen(print, onDone: (){print(' 完成')})。 ", the "完成 :"never work
从 Redis 获取消息时，onDone:(){print('done')} 从未起作用。 import 'package:dartis/dartis.dart' as redis show PubS
Vim状态栏的预测/完成？
昨天我玩了一些vim脚本，并设法通过循环来对当前输入的内容进行状态栏预测(请参见屏幕截图(灰色+黄色栏))。问题是，我不记得我是怎么得到的，也找不到我用于该vim魔术的代码片段(我记得它很简单):它
Bash 完成
我尝试加载 bash_completion在我的 bash (3.2.25) 中，它不起作用。没有消息等。我在我的 .bashrc 中使用了以下内容 if [ -f ~/.bash_completio
具有等号和可枚举标志值的 Bash 完成
我正在尝试构建一个 bash 完成例程，它将建议命令行标志和合适的标志值。例如在下面 fstcompose 命令我想比赛套路先建议 compose_filter= 标志，然后建议来自 [alt_seq
重定向符号后的 Bash 完成
当我尝试在重定向符号后完成路径时，bash 完成的行为就好像它仍在尝试在重定向之前完成命令的参数一样。例如: dpkg -l > /med标签通过在 /med 之后点击 Tab我希望它完成通往 /
iphone - CAKeyframeAnimation 完成
我的类中有几个 CAKeyframeAnimation 对象。他们都以 self 为代表。在我的animationDidStop函数中，我如何知道调用来自哪里？是否有任何变量可以传递给 CAKe
cocoa - NSDateFormatter 完成
我有一个带有 NSDateFormatter 的 NSTextField。格式化程序接受“mm/dd/yy”。可以自动补全日期吗？因此，用户可以输入“mm”，格式化程序将完成当前月份和年份。最佳答
cocoa - NSTextfield 完成
有一个解决方案可以使用以下方法完成 NSTextField : - (NSArray *)control:(NSControl *)control textView:(NSTextView *)tex
javascript - 完成()与返回完成()
我正在阅读 Passport 的文档，我注意到 serialize()和 deserialize() done()被调用而不被返回。但是，当使用 passport.use() 设置新策略时在回调函数
javascript 加载图像!完成
在 ubuntu 11.10 上的 Firefox 8.0 中，尽管 img.complete 为 false，但仍会调用 onload 函数 draw。我设法用 setTimeout hack 解决
c++ - 等待第一个将来用C++完成
假设我有两个与两个并行执行的计算相对应的 future 。我如何等到第一个 future 准备好？理想情况下，我正在寻找类似于Python asyncio's wait且参数为return_when=
Java 数据结构表明队列已结束/完成？
我正在寻找一种 Java 7 数据结构，其行为类似于 java.util.Queue，并且还具有“最终项目已被删除”的概念。例如，应可以表达如下概念: while(!endingQueue.isFi
jquery - 完成 If 语句
这是一个简单的问题。 if ($('.dataTablePageList')) { 我想做的是执行一个 if 语句，该语句表示如果具有 dataTablesPageList 类的对象也具有 menu
jQuery 在执行之前等待replaceWith 完成
我用replaceWith批量替换了许多div中的html。替换后，我使用 jTruncate 来截断文本。然而它不起作用，因为在执行时，replaceWith 还没有完成。我尝试了回调技巧 ( H
JavaScript 表单提交()完成
有没有办法调用 javascript 表单 submit() 函数或 JQuery $.submit() 函数并确保它完成提交过程？具体来说，在一个表单中，我试图在一个 IFrame 中提交一个表单。
javascript - 推迟行动直到 .each() 完成
我有以下方法: function animatePortfolio(fadeElement) { fadeElement.children('article').each(function(i
android - registerEntityModifier 完成
我刚刚开始使用 AndEngine，我正在像这样移动 Sprite : if(pValueY < 0 && !jumping) { jumping =
android - 完成 "all"异步任务后更新屏幕
我正在使用 asynctask 来执行冗长的操作，例如数据库读取。我想开始一个新 Activity 并在所有异步任务完成后呈现其内容。实现这一目标的最佳方法是什么？我知道 onPostExecute
从另一个完成 Bash 完成
我有一个脚本需要命令名称和该命令的参数作为参数。所以我想编写一个完成函数来完成命令的名称并完成该命令的参数。所以我可以这样完成命令的名称 if [[ "$COMP_CWORD" == 1 ]];
android - 完成()不工作
我的应用程序有一个相当奇怪的行为。我在 BOOT_COMPLETE 之后启动我的应用程序，因此在我启动设备后它是可见的。 GUI 响应迅速，一切正常，直到我调用 finish()，按下按钮时，什么都没

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

django - django-haystack自动完成返回的结果太宽