elasticsearch - 如何从elasticsearch 6.1搜索中排除字段？-6ren

elasticsearch - 如何从elasticsearch 6.1搜索中排除字段？

转载作者：行者123 更新时间：2023-12-01 00:32:17

我有一个包含多个字段的索引。我想根据所有字段中搜索字符串的存在进行过滤，除了一个 - 用户评论 .
我正在做的查询搜索是

{
    "from": offset,
    "size": limit,
    "_source": [
      "document_title"
    ],
    "query": {
      "function_score": {
        "query": {
          "bool": {
            "must":
            {
              "query_string": {
                "query": "#{query}"
              }
            }
          }
        }
      }
    }
  }

尽管查询字符串正在搜索所有字段，并在 中为我提供了具有匹配字符串的文档用户评论场也是如此。但是，我想针对所有不包含 的字段进行查询。用户评论 field 。
白名单是一个非常大的列表，而且字段的名称是动态的，因此使用 fields 参数提及白名单字段列表是不可行的。

"query_string": {
                    "query": "#{query}",
                    "fields": [
                      "document_title",
                      "field2"
                    ]
                  }

任何人都可以提出一个关于如何从搜索中排除字段的想法吗？

最佳答案

有一种方法可以让它工作，它不是很漂亮，但可以完成工作。您可以使用 boost 实现您的目标。和 multifield query_string的参数, bool 查询以结合分数和设置 min_score :

POST my-query-string/doc/_search
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
            "query": "#{query}",
            "type": "most_fields",
            "boost": 1
          }
        },
        {
          "query_string": {
            "fields": [
              "comments"
            ],
            "query": "#{query}",
            "boost": -1
          }
        }
      ]
    }
  },
  "min_score": 0.00001
}

那么引擎盖下会发生什么？
假设您有以下一组文档:

PUT my-query-string/doc/1
{
  "title": "Prodigy in Bristol",
  "text": "Prodigy in Bristol",
  "comments": "Prodigy in Bristol"
}
PUT my-query-string/doc/2
{
  "title": "Prodigy in Birmigham",
  "text": "Prodigy in Birmigham",
  "comments": "And also in Bristol"
}
PUT my-query-string/doc/3
{
  "title": "Prodigy in Birmigham",
  "text": "Prodigy in Birmigham and Bristol",
  "comments": "And also in Cardiff"
}
PUT my-query-string/doc/4
{
  "title": "Prodigy in Birmigham",
  "text": "Prodigy in Birmigham",
  "comments": "And also in Cardiff"
}

在您的搜索请求中，您只想查看文档 1 和 3，但您的原始查询将返回 1、2 和 3。
在 Elasticsearch 中，搜索结果按 relevance _score 排序，分数越大越好。
所以让我们尝试 boost下 "comments"字段，因此它对相关性分数的影响被忽略。我们可以通过将两个查询与 should 结合来实现这一点。并使用负值 boost :

POST my-query-string/doc/_search
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
            "query": "Bristol"
          }
        },
        {
          "query_string": {
            "fields": [
              "comments"
            ],
            "query": "Bristol",
            "boost": -1
          }
        }
      ]
    }
  }
}

这将为我们提供以下输出:

{
  "hits": {
    "total": 3,
    "max_score": 0.2876821,
    "hits": [
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "3",
        "_score": 0.2876821,
        "_source": {
          "title": "Prodigy in Birmigham",
          "text": "Prodigy in Birmigham and Bristol",
          "comments": "And also in Cardiff"
        }
      },
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "2",
        "_score": 0,
        "_source": {
          "title": "Prodigy in Birmigham",
          "text": "Prodigy in Birmigham",
          "comments": "And also in Bristol"
        }
      },
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "1",
        "_score": 0,
        "_source": {
          "title": "Prodigy in Bristol",
          "text": "Prodigy in Bristol",
          "comments": "Prodigy in Bristol",
          "discount_percent": 10
        }
      }
    ]
  }
}

文档 2 受到了惩罚，但文档 1 也受到了惩罚，尽管它是我们想要的匹配项。为什么会这样？
以下是 Elasticsearch 的计算方式 _score在这种情况下:

_score = max(title:"Bristol", text:"Bristol", comments:"Bristol") - comments:"Bristol"

文档 1 匹配 comments:"Bristol"部分，它也恰好是最好的分数。根据我们的公式，结果分数为 0。
我们实际上想要做的是提升第一个子句(带有“所有”字段) 更多如果匹配更多字段。
我们可以提升吗 query_string匹配更多字段？
我们可以， query_string在 multifield模式有一个 type参数就是这样做的。查询将如下所示:

POST my-query-string/doc/_search
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
            "type": "most_fields",
            "query": "Bristol"
          }
        },
        {
          "query_string": {
            "fields": [
              "comments"
            ],
            "query": "Bristol",
            "boost": -1
          }
        }
      ]
    }
  }
}

这将为我们提供以下输出:

{
  "hits": {
    "total": 3,
    "max_score": 0.57536423,
    "hits": [
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "1",
        "_score": 0.57536423,
        "_source": {
          "title": "Prodigy in Bristol",
          "text": "Prodigy in Bristol",
          "comments": "Prodigy in Bristol",
          "discount_percent": 10
        }
      },
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "3",
        "_score": 0.2876821,
        "_source": {
          "title": "Prodigy in Birmigham",
          "text": "Prodigy in Birmigham and Bristol",
          "comments": "And also in Cardiff"
        }
      },
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "2",
        "_score": 0,
        "_source": {
          "title": "Prodigy in Birmigham",
          "text": "Prodigy in Birmigham",
          "comments": "And also in Bristol"
        }
      }
    ]
  }
}

如您所见，不需要的文档 2 位于底部，其得分为 0。这是这次得分的计算方式:

_score = sum(title:"Bristol", text:"Bristol", comments:"Bristol") - comments:"Bristol"

所以文档匹配 "Bristol"在任何领域都被选中。 comments:"Bristol" 的相关性得分被淘汰了，只有匹配的文档 title:"Bristol"或 text:"Bristol"得到了 _score > 0。
我们可以过滤掉那些分数不理想的结果吗？
是的，我们可以，使用 min_score :

POST my-query-string/doc/_search
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
            "query": "Bristol",
            "type": "most_fields",
            "boost": 1
          }
        },
        {
          "query_string": {
            "fields": [
              "comments"
            ],
            "query": "Bristol",
            "boost": -1
          }
        }
      ]
    }
  },
  "min_score": 0.00001
}

这将起作用(在我们的例子中)，因为当且仅当 "Bristol" 文档的分数将为 0匹配字段 "comments" only 并且不匹配任何其他字段。
输出将是:

{
  "hits": {
    "total": 2,
    "max_score": 0.57536423,
    "hits": [
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "1",
        "_score": 0.57536423,
        "_source": {
          "title": "Prodigy in Bristol",
          "text": "Prodigy in Bristol",
          "comments": "Prodigy in Bristol",
          "discount_percent": 10
        }
      },
      {
        "_index": "my-query-string",
        "_type": "doc",
        "_id": "3",
        "_score": 0.2876821,
        "_source": {
          "title": "Prodigy in Birmigham",
          "text": "Prodigy in Birmigham and Bristol",
          "comments": "And also in Cardiff"
        }
      }
    ]
  }
}

可以以不同的方式完成吗？
当然。我实际上不建议使用 _score调整，因为这是一个非常复杂的问题。
我建议获取现有映射并构建一个字段列表以预先运行查询，这将使代码更加简单明了。
答案中提出的原始解决方案(保留历史记录)
最初建议使用这种查询，其意图与上述解决方案完全相同:

POST my-query-string/doc/_search
{
  "query": {
    "function_score": {
      "query": {
        "bool": {
          "must": {
            "query_string": {
              "fields" : ["*", "comments^0"],
              "query": "#{query}"
            }
          }
        }
      }
    }
  },
  "min_score": 0.00001
}

唯一的问题是，如果索引包含任何数值，这部分:

"fields": ["*"]

引发错误，因为文本查询字符串不能应用于数字。

关于elasticsearch - 如何从elasticsearch 6.1搜索中排除字段？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/52757277/

文章推荐： Python:scipy/numpy 两个一维向量之间的所有对计算

文章推荐： xcode - MonoDevelop:安装到 SSD 驱动器以提高速度

文章推荐： php - 计算PHP中两次之间的差异(以小时为单位)

mysql修改记录时update操作字段=字段+字符串
在有些场景下，我们需要对我们的varchar类型的字段做修改，而修改的结果为两个字段的拼接或者一个字段+字符串的拼接。如下所示，我们希望将xx_role表中的name修改为name+id。
MySQL SUM IF 字段 b = 字段 a
SELECT incMonth as Month, SUM( IF(item_type IN('typ1', 'typ2') AND incMonth = Month, 1, 0 ) )AS
java - 如果直接从内存读取 volatile 字段，那么从哪里读取非 volatile 字段？
我最近读到 volatile 字段是线程安全的，因为 When we use volatile keyword with a variable, all the threads read its va
python - 在数据库中已有数据之后添加的 UUID 字段。有没有办法为现有数据填充 UUID 字段？
我在一些模型中添加了一个 UUID 字段，然后使用 South 进行了迁移。我创建的任何新对象都正确填充了 UUID 字段。但是，我所有旧数据的 UUID 字段为空。有没有办法为现有数据填充 UUI
php - 左连接中的两个表都有 id 字段。尝试从第一个数据库中提取 id 字段，但获取第二个数据库
刚刚将我的网站从 mysql_ 更新为 mysqli，并破坏了之前正常运行的查询。我试图从旋转中提取 id，因为它每次都会增加 1，但我不断获取玩家 id，有人可以告诉我我做错了什么吗？我尝试了将
mysql - 如何使用 MySQL 将一个表中的一列(字段)复制到另一个表的空列(字段)，这两个表都是同一数据库的一部分？
我在 Mac OS X 上使用带有 Sequel Pro 的 MySQL。我想将一个表中的一个字段(即名为“GAME_DY”的列)复制到另一个名为“DAY_ID”的表的空字段中。两个表都是同一数据库的
java - 为序列化设置一个 transient 字段，但为 JPA 设置非 transient 字段
问题: 是否有可能有一个字段被 JPA 保留但被序列化跳过？可以实现相反的效果(JPA 跳过字段而序列化则不会)，如果使用此功能，那么相反的操作肯定会很有用。类似这样的事情: @Entity cl
php - 无重复(分组依据)字段 1 循环，字段 2 位于水平线
假设我有一个名为“dp”的表 Year | Month | Payment| Payer_ID | Payment_Recipient | 2008/2009 | July
c - 我在 IP header 中找不到 DSCP 字段，只有已弃用的 TOS 字段
我将尝试通过我的 Raspberry Pi 接入点保证一些 QoS。开始之前，我先动手:我阅读了有关 tcp、udp 和 ip header 的内容。在IP header description我看
dart - 什么时候应该在 dart 中使用 final 字段、工厂构造函数或带有 getter 的私有(private)字段？
如果你能弄清楚如何重命名这个问题，我愿意接受建议。在 Dart 语言中，可以编写一个带有 final 字段的类。这些是只能设置的字段构造函数前 body 跑。这可以在声明中(通常用于类中的静态常量)
javascript - jquery:使用两个带有两个字段(字段 1、字段 2 + 1 天)的日期选择器，例如 booking.com
你怎么样？我有两个带有两个字段的日期选择器我希望当用户选择 (From) 时，第二个字段 (TO) 将是 next day 。比如 booking.com 例如:当用户选择From 01-01-2
mysql - 将字段从 T1 字段 A 复制到 T2 字段 A where(if or when) T1 field B = T2 field B (mysql)
我想我已经看到了这个问题的一些答案，这些答案可能与我需要的相差不远，但我对 mysql 的了解还不够确定，所以我会根据我的具体情况提出问题。我有一个包含多个表的数据库，为此，如果“image”表上的
mySQL在单个查询中多次使用相同的表/字段
我在 mySQL 数据库中有 2 个表: customers ============ customer_id (1, 2 ) customer_name (john, mark) orders ==
数据库归档与基于时间段的表/字段
我正在开发一个员工目标 Web 应用程序。领导/经理在与团队成员讨论后为他们设定目标。这是一年/半年/季度，具体取决于组织遵循的评估周期。现在的问题是添加基于时间段的字段或存档上一季度/年度数据的
Sitecore 字段，用于从媒体库中选择多个文件并能够上传文件
我正在寻找允许内容编辑器从媒体库中选择多个文件的东西，这些文件将在渲染中列出。他们还需要能够上传文件和搜索。它必须在页面编辑器(版本 8 中称为体验编辑器)中工作。到目前为止我所考虑的: 一堆文件字
r - 创建 "other"字段
现在，我有以下由 original.df %.% group_by(Category) %.% tally() %.% arrange(desc(n)) 创建的 data.frame。 DF 5),
潘塔霍。将登录的错误消息放入字符串/字段
我想知道是否有一些步骤/解决方案可以处理错误消息并将它们放入 Pentaho 工具中的某个字符串或字段中？例如，如果连接到数据库时发生某些错误，则将该消息从登录到字符串/字段。最佳答案我们在作业的
iPhone如何制作 "To"字段，如短信应用程序
如何制作像短信应用程序一样的“收件人”字段？例如，右侧有一个“+”按钮，当添加某人时，名称将突出显示并可单击，如圆角矩形等。有没有内置的框架？最佳答案不，但请参阅 Three20 的 TTMess
delphi - 列出记录的元素\字段
是否可以获取记录的元素或字段的列表通过类型信息类似于类的已发布属性的列表吗？谢谢！最佳答案取决于您的delphi版本，如果您使用的是delphi 2010或更高版本，则可以使用“新rtti”
带有外键列表的 SQLite 字段
我正在构建一个 SQLite 数据库来保存我的房地产经纪人的列表。我已经能够使用外键来识别每个代理的列表，但我想在每个代理的记录中创建一个列表；从代理商和列表之间的一对一关系转变为一对多关系。看这里

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

elasticsearch - 如何从elasticsearch 6.1搜索中排除字段？