java - 如何编写用户定义的聚合函数？-6ren

java - 如何编写用户定义的聚合函数？

转载作者：行者123 更新时间：2023-11-29 07:29:33

我正在尝试理解 Java Spark 文档。有一个名为Untyped User Defined Aggregate Functions 的部分，其中包含一些我无法理解的示例代码。这是代码:

package org.apache.spark.examples.sql;

// $example on:untyped_custom_aggregation$
import java.util.ArrayList;
import java.util.List;

import org.apache.spark.sql.Dataset;
import org.apache.spark.sql.Row;
import org.apache.spark.sql.SparkSession;
import org.apache.spark.sql.expressions.MutableAggregationBuffer;
import org.apache.spark.sql.expressions.UserDefinedAggregateFunction;
import org.apache.spark.sql.types.DataType;
import org.apache.spark.sql.types.DataTypes;
import org.apache.spark.sql.types.StructField;
import org.apache.spark.sql.types.StructType;
// $example off:untyped_custom_aggregation$

public class JavaUserDefinedUntypedAggregation {

  // $example on:untyped_custom_aggregation$
  public static class MyAverage extends UserDefinedAggregateFunction {

    private StructType inputSchema;
    private StructType bufferSchema;

    public MyAverage() {
      List<StructField> inputFields = new ArrayList<>();
      inputFields.add(DataTypes.createStructField("inputColumn", DataTypes.LongType, true));
      inputSchema = DataTypes.createStructType(inputFields);

      List<StructField> bufferFields = new ArrayList<>();
      bufferFields.add(DataTypes.createStructField("sum", DataTypes.LongType, true));
      bufferFields.add(DataTypes.createStructField("count", DataTypes.LongType, true));
      bufferSchema = DataTypes.createStructType(bufferFields);
    }
    // Data types of input arguments of this aggregate function
    public StructType inputSchema() {
      return inputSchema;
    }
    // Data types of values in the aggregation buffer
    public StructType bufferSchema() {
      return bufferSchema;
    }
    // The data type of the returned value
    public DataType dataType() {
      return DataTypes.DoubleType;
    }
    // Whether this function always returns the same output on the identical input
    public boolean deterministic() {
      return true;
    }
    // Initializes the given aggregation buffer. The buffer itself is a `Row` that in addition to
    // standard methods like retrieving a value at an index (e.g., get(), getBoolean()), provides
    // the opportunity to update its values. Note that arrays and maps inside the buffer are still
    // immutable.
    public void initialize(MutableAggregationBuffer buffer) {
      buffer.update(0, 0L);
      buffer.update(1, 0L);
    }
    // Updates the given aggregation buffer `buffer` with new input data from `input`
    public void update(MutableAggregationBuffer buffer, Row input) {
      if (!input.isNullAt(0)) {
        long updatedSum = buffer.getLong(0) + input.getLong(0);
        long updatedCount = buffer.getLong(1) + 1;
        buffer.update(0, updatedSum);
        buffer.update(1, updatedCount);
      }
    }
    // Merges two aggregation buffers and stores the updated buffer values back to `buffer1`
    public void merge(MutableAggregationBuffer buffer1, Row buffer2) {
      long mergedSum = buffer1.getLong(0) + buffer2.getLong(0);
      long mergedCount = buffer1.getLong(1) + buffer2.getLong(1);
      buffer1.update(0, mergedSum);
      buffer1.update(1, mergedCount);
    }
    // Calculates the final result
    public Double evaluate(Row buffer) {
      return ((double) buffer.getLong(0)) / buffer.getLong(1);
    }
  }
  // $example off:untyped_custom_aggregation$

  public static void main(String[] args) {
    SparkSession spark = SparkSession
      .builder()
      .appName("Java Spark SQL user-defined DataFrames aggregation example")
      .getOrCreate();

    // $example on:untyped_custom_aggregation$
    // Register the function to access it
    spark.udf().register("myAverage", new MyAverage());

    Dataset<Row> df = spark.read().json("examples/src/main/resources/employees.json");
    df.createOrReplaceTempView("employees");
    df.show();
    // +-------+------+
    // |   name|salary|
    // +-------+------+
    // |Michael|  3000|
    // |   Andy|  4500|
    // | Justin|  3500|
    // |  Berta|  4000|
    // +-------+------+

    Dataset<Row> result = spark.sql("SELECT myAverage(salary) as average_salary FROM employees");
    result.show();
    // +--------------+
    // |average_salary|
    // +--------------+
    // |        3750.0|
    // +--------------+
    // $example off:untyped_custom_aggregation$

    spark.stop();
  }
}

我对上述代码的疑惑是:

每当我想创建一个 UDF 时，我是否应该拥有函数 initialize、update 和 merge？
变量inputSchema 和bufferSchema 有什么意义？我很惊讶它们的存在，因为它们从来没有被用来创建任何 DataFrame。它们应该出现在每个 UDF 中吗？如果是，那么它们应该是完全相同的名字吗？
为什么 inputSchema 和 bufferSchema 的 getter 没有命名为 getInputSchema() 和 getBufferSchema()？为什么没有这些变量的 setter ？
这里称为deterministic() 的函数有什么意义？请给出调用此函数有用的场景。

总的来说，我想知道如何在 Spark 中编写用户定义的聚合函数。

最佳答案

Whenever I want to create a UDF, should I have the functions initialize, update and merge

UDF 代表用户定义的函数，而方法initialize、update 和merge 用于用户定义的聚合函数(又名UDAF)。

UDF 是一个函数，它处理单行以(通常)生成一行(例如 upper 函数)。

UDAF 是一种使用零行或多行生成一行的函数(例如，count 聚合函数)。

您当然不必(也不可能)为用户提供函数initialize、update 和merge -定义函数 (UDF)。

使用任何 udf functions定义和注册 UDF。

val myUpper = udf { (s: String) => s.toUpperCase }

How to how to write a user defined aggregate function in Spark.

What is the significance of the variables inputSchema and bufferSchema?

(无耻插件:我一直在 UserDefinedAggregateFunction — Contract for User-Defined Untyped Aggregate Functions (UDAFs) 的 Mastering Spark SQL 一书中描述 UDAF)

引用 Untyped User-Defined Aggregate Functions :

// Data types of input arguments of this aggregate function
def inputSchema: StructType = StructType(StructField("inputColumn", LongType) :: Nil)

// Data types of values in the aggregation buffer
def bufferSchema: StructType = {
  StructType(StructField("sum", LongType) :: StructField("count", LongType) :: Nil)
}

换句话说，inputSchema 是您对输入的期望，而 bufferSchema 是您在进行聚合时临时保留的内容。

Why are there no setters of these variables?

它们是由 Spark 管理的扩展点。

What is the significance of the function called deterministic() here?

引用 Untyped User-Defined Aggregate Functions :

// Whether this function always returns the same output on the identical input
def deterministic: Boolean = true
Please give a scenario when it would be useful to call this function.

这是我仍在努力的事情，所以今天无法回答。

关于java - 如何编写用户定义的聚合函数？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/44936394/

文章推荐： mysql - 如何访问 MySQL 过程参数？

文章推荐： java - 为什么使用 PosixFilePermission 设置目录权限不起作用

文章推荐： php - MySQL - 从多列中选择并忽略重复结果

c - 如何从客户端(用 C 编写)接收 int 数组到服务器(用 python 编写)
我只想从客户端向服务器发送数组 adc_array=[w, x, y, z]。下面是客户端代码，而我的服务器是在只接受 json 的 python 中。编译代码时我没有收到任何错误，但收到 2 条警告
node.js - 如何连接我的移动应用程序(用 lua 编写)和我的服务器(用 node.js 编写)？
我是 lua 和 Node js 的新手，我正在尝试将我正在开发的移动应用程序连接到服务器。问题是它连接到服务器，但我尝试传递的数据丢失或无法到达服务器。对我正在做的事情有什么问题有什么想法吗？ th
Haskell 编写 myLength
我在这个页面上工作 http://www.haskell.org/haskellwiki/99_questions/Solutions/4 我理解每个函数的含义，看到一个函数可以像这样以多种方式定义，
Java CSV 编写
我目前正在尝试将数据写入 excel 以生成报告。我可以将数据写入 csv 文件，但它不会按照我想要的顺序出现在 excel 中。我需要数据在每列的最佳和最差适应性下打印，而不是全部打印在平均值下。这
Java - 编写、读取和修改带参数的字符串
所以，我正在做一个项目，现在我有一个问题，所以我想得到你的帮助:) 首先，我已经知道如何编写和读取 .txt 文件，但我想要的不仅仅是 x.hasNext()。我想知道如何像 .ini 那样编写、读
javascript - 编写 For 循环来计算阶乘
我正在尝试编写一个函数，该函数将返回作为输入给出的任何数字的阶乘。现在，我的代码绝对是一团糟。请帮忙。 function factorialize(num) { for (var i=num, i
Javascript，编写 if 条件的更好方法
这个问题已经有答案了: Check variable equality against a list of values (16 个回答) 已关闭 4 年前。有没有一种简洁或更好的方法来编写这个条件
aframe - 编写 A 型框架的测试规范
我对 VR 完全陌生，正在 AFrame 中为一个类(class)项目开发 VR 太空射击游戏，并且想知道 AFrame 中是否有 TDD 的任何文档/标准。有人能指出我正确的方向吗？最佳答案几乎
javascript - 编写 for 循环以使用数组创建多个方法
我正在尝试创建一个 for 循环，它将重现以下功能代码块，但以一种更具吸引力的方式。这是与 Soundcould 小部件 API 实现一起使用的 here on stackoverflow $(doc
Java 编写/编辑属性文件
我有一个非常令人困惑的问题。我正在尝试更改属性文件中的属性，但它只是没有更改... 这是代码: package config; import java.io.FileNotFoundException
aframe - 编写 A 型框架的测试规范
我对 VR 完全陌生，正在 AFrame 中为一个类(class)项目开发 VR 太空射击游戏，并且想知道 AFrame 中是否有 TDD 的任何文档/标准。有人能指出我正确的方向吗？最佳答案几乎
.net - 编写.NET互操作调试器
我正在开发一个用户模式(Ring3)代码级调试器。它还应支持.NET可执行文件的本机(x86)调试。基本上，我需要执行以下操作: 1).NET在隐身模式下加载某些模块，而没有LOAD_DLL_DEBU
python - 编写 if 语句以避免某些列表项的更好方法是什么？
我有一个列表，我知道有些项目是不必要打印的，我正在尝试通过 if 语句来做到这一点...但是它变得非常复杂，所以有没有什么方法可以在 if 语句中包含多个索引而无需打印重写整个声明。看起来像这样的东
c# - 编写 if 语句是否会以不同方式影响程序的速度和效率？
我很好奇以不同方式编写 if 语句是否会影响程序的速度和效率。所以，例如写一个这样的: bool isActive = true; bool isResponding = false; if (isA
javascript - 编写 if 语句的新方法
我在搜索网站的源代码时找到了一种以另一种方式(我认为)编写 if 语句的方法。代替: if(a)b; 或: a?b:''; 我读了: !a||b; 第三种方式和前两种方式一样吗？如果是，为什么我们要
Java + 编写 XML
我的数据采用以下格式(HashMap的列表) {TeamName=India, Name=Sachin, Score=170} {TeamName=India, Name=Sehwag, Score=
mysql - 编写 HAVING 条件的最有效方法
我目前正在完成 More JOIN operations sqlzoo 的教程，遇到了下面的代码作为#12 的答案: SELECT yr,COUNT(title) FROM movie JOIN ca
ruby - 编写 && 检查列表的更好方法？
我正试图找到一种更好的方法来编写这段代码: def down_up(array, player) 7.downto(3).each do |row| 8.times do |col
由 C++ 编写
出于某种原因，我的缓冲区中充满了乱码，我不确定为什么。我什至用十六进制编辑器检查了我的文件，以验证我的字符是否以 2 字节的 unicode 格式保存。我不确定出了什么问题。 [打开文件] fseek
c# - 编写 FizzBuzz
阅读编码恐怖片时，我刚刚又遇到了 FizzBuzz。原帖在这里:Coding Horror: Why Can't Programmers.. Program? 对于那些不知道的人:FizzBu

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

java - 如何编写用户定义的聚合函数？