gpt4 book ai didi

Elasticsearch phrase suggester 向我建议我的索引中不存在的建议

转载 作者:行者123 更新时间:2023-11-29 02:53:49 25 4
gpt4 key购买 nike

我有一个 Elasticsearch 索引,其中包含一些数据。我实现了 did-you-mean 功能,因此当用户写错拼写的内容时,它可以收到包含正确单词的建议。

我使用短语 suggester 是因为我需要对短语的建议,例如名称,问题是索引中不存在某些建议。

例子:

document in the index: coding like a master
search: Codning like a boss
suggestion: <em>coding</em> like a boss
search result: not found

我的问题是,我的索引中没有匹配指定建议的短语,所以它向我推荐不存在的短语,因此会给我一个未找到的搜索。

我能用它做什么? phrase suggester 不应该为索引中实际存在的短语提供建议吗?

这里我会留下相应的查询、映射和设置,以备不时之需。

设置和映射

{
"settings": {
"index": {
"number_of_shards": 3,
"number_of_replicas": 1,
"search.slowlog.threshold.fetch.warn": "2s",
"index.analysis.analyzer.default.filter.0": "standard",
"index.analysis.analyzer.default.tokenizer": "standard",
"index.analysis.analyzer.default.filter.1": "lowercase",
"index.analysis.analyzer.default.filter.2": "asciifolding",
"index.priority": 3,
"analysis": {
"analyzer": {
"suggests_analyzer": {
"tokenizer": "lowercase",
"filter": [
"lowercase",
"asciifolding",
"shingle_filter"
],
"type": "custom"
}
},
"filter": {
"shingle_filter": {
"min_shingle_size": 2,
"max_shingle_size": 3,
"type": "shingle"
}
}
}
}
},
"mappings": {
"my_type": {
"properties": {
"suggest_field": {
"analyzer": "suggests_analyzer",
"type": "string"
}
}
}
}
}

查询

{
"DidYouMean": {
"text": "Codning like a boss",
"phrase": {
"field": "suggest_field",
"size": 1,
"gram_size": 1,
"confidence": 2.0
}
}
}

感谢您的帮助。

最佳答案

这实际上是意料之中的。如果您使用 analyze api 分析文档,您将更好地了解正在发生的事情。

GET suggest_index/_analyze?text=coding like a master&analyzer=suggests_analyzer

这是输出

{
"tokens": [
{
"token": "coding",
"start_offset": 0,
"end_offset": 6,
"type": "word",
"position": 1
},
{
"token": "coding like",
"start_offset": 0,
"end_offset": 11,
"type": "shingle",
"position": 1
},
{
"token": "coding like a",
"start_offset": 0,
"end_offset": 13,
"type": "shingle",
"position": 1
},
{
"token": "like",
"start_offset": 7,
"end_offset": 11,
"type": "word",
"position": 2
},
{
"token": "like a",
"start_offset": 7,
"end_offset": 13,
"type": "shingle",
"position": 2
},
{
"token": "like a master",
"start_offset": 7,
"end_offset": 20,
"type": "shingle",
"position": 2
},
{
"token": "a",
"start_offset": 12,
"end_offset": 13,
"type": "word",
"position": 3
},
{
"token": "a master",
"start_offset": 12,
"end_offset": 20,
"type": "shingle",
"position": 3
},
{
"token": "master",
"start_offset": 14,
"end_offset": 20,
"type": "word",
"position": 4
}
]
}

如您所见,为文本生成了一个标记“编码”,因此它在您的索引中。它向您建议不在索引中的内容。如果您严格想要短语搜索,那么您可能需要考虑使用 keyword tokenizer .例如,如果您将映射更改为类似

{
"settings": {
"index": {
"analysis": {
"analyzer": {
"suggests_analyzer": {
"tokenizer": "lowercase",
"filter": [
"lowercase",
"asciifolding",
"shingle_filter"
],
"type": "custom"
},
"raw_analyzer": {
"tokenizer": "keyword",
"filter": [
"lowercase",
"asciifolding"
]
}
},
"filter": {
"shingle_filter": {
"min_shingle_size": 2,
"max_shingle_size": 3,
"type": "shingle"
}
}
}
}
},
"mappings": {
"my_type": {
"properties": {
"suggest_field": {
"analyzer": "suggests_analyzer",
"type": "string",
"fields": {
"raw": {
"analyzer": "raw_analyzer",
"type": "string"
}
}
}
}
}
}
}

那么这个查询会给你预期的结果

{
"DidYouMean": {
"text": "codning lke a master",
"phrase": {
"field": "suggest_field.raw",
"size": 1,
"gram_size": 1
}
}
}

它不会显示任何“像老板一样编码”

编辑 1

2) 根据您的评论以及在我自己的数据集上运行一些短语建议,我觉得更好的方法是使用 collate选项 phrase suggester 提供这样我们就可以根据 query 检查每个建议,并且仅当它要从索引中取回任何文档时才返回建议。我还在映射中添加了 词干分析器 以仅考虑词根。我正在使用 light_english 因为它不那么激进。 More在那上面。

映射的分析器部分现在看起来像这样

 "analysis": {
"analyzer": {
"suggests_analyzer": {
"tokenizer": "standard",
"filter": [
"lowercase",
"english_possessive_stemmer",
"light_english_stemmer",
"asciifolding",
"shingle_filter"
],
"type": "custom"
}
},
"filter": {
"light_english_stemmer": {
"type": "stemmer",
"language": "light_english"
},
"english_possessive_stemmer": {
"type": "stemmer",
"language": "possessive_english"
},
"shingle_filter": {
"min_shingle_size": 2,
"max_shingle_size": 4,
"type": "shingle"
}
}
}

现在这个查询会给你想要的结果。

{
"suggest" : {
"text" : "appel on the tabel",
"simple_phrase" : {
"phrase" : {
"field" : "suggest_field",
"size" : 5,
"collate": {
"query": {
"inline" : {
"match_phrase": {
"{{field_name}}" : "{{suggestion}}"
}
}
},
"params": {"field_name" : "suggest_field"},
"prune": false
}
}
}
},
"size": 0
}

这会让你回到 table 上的苹果这里使用了 match_phrase 查询,它将根据索引运行每个建议的短语。您可以设置 "prune": true 并查看所有建议的结果,而不管是否匹配。您可能需要考虑使用 stop 过滤器来避免停用词。

希望这对您有所帮助!

关于Elasticsearch phrase suggester 向我建议我的索引中不存在的建议,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/34819812/

25 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com