python - 在 PySpark 中使用 Apache Spark 数据帧删除重音的最佳方法是什么？-6ren

python - 在 PySpark 中使用 Apache Spark 数据帧删除重音的最佳方法是什么？

转载作者：太空狗更新时间：2023-10-29 17:19:28

我需要从不同数据集中删除西类牙语和其他语言字符的重音。

我已经根据此 post 中提供的代码做了一个函数删除特殊的口音。问题在于该函数运行缓慢，因为它使用了 UDF。我只是想知道我是否可以提高函数的性能以在更短的时间内获得结果，因为这对小数据帧有好处，但对大数据帧不利。

提前致谢。

这里是代码，您将能够按照显示的方式运行它:

# Importing sql types
from pyspark.sql.types import StringType, IntegerType, StructType, StructField
from pyspark.sql.functions import udf, col
import unicodedata

# Building a simple dataframe:
schema = StructType([StructField("city", StringType(), True),
                     StructField("country", StringType(), True),
                     StructField("population", IntegerType(), True)])

countries = ['Venezuela', 'US@A', 'Brazil', 'Spain']
cities = ['Maracaibó', 'New York', '   São Paulo   ', '~Madrid']
population = [37800000,19795791,12341418,6489162]

# Dataframe:
df = sqlContext.createDataFrame(list(zip(cities, countries, population)), schema=schema)

df.show()

class Test():
    def __init__(self, df):
        self.df = df

    def clearAccents(self, columns):
        """This function deletes accents in strings column dataFrames, 
        it does not eliminate main characters, but only deletes special tildes.

        :param columns  String or a list of column names.
        """
        # Filters all string columns in dataFrame
        validCols = [c for (c, t) in filter(lambda t: t[1] == 'string', self.df.dtypes)]

        # If None or [] is provided with column parameter:
        if (columns == "*"): columns = validCols[:]

        # Receives  a string as an argument
        def remove_accents(inputStr):
            # first, normalize strings:
            nfkdStr = unicodedata.normalize('NFKD', inputStr)
            # Keep chars that has no other char combined (i.e. accents chars)
            withOutAccents = u"".join([c for c in nfkdStr if not unicodedata.combining(c)])
            return withOutAccents

        function = udf(lambda x: remove_accents(x) if x != None else x, StringType())
        exprs = [function(col(c)).alias(c) if (c in columns) and (c in validCols) else c for c in self.df.columns]
        self.df = self.df.select(*exprs)

foo = Test(df)
foo.clearAccents(columns="*")
foo.df.show()

最佳答案

一个可能的改进是构建自定义 Transformer ，它将处理 Unicode 规范化和相应的 Python 包装器。它应该减少在 JVM 和 Python 之间传递数据的总体开销，并且不需要对 Spark 本身进行任何修改或访问私有(private) API。

在 JVM 端你需要一个类似于这个的转换器:

package net.zero323.spark.ml.feature

import java.text.Normalizer
import org.apache.spark.ml.UnaryTransformer
import org.apache.spark.ml.param._
import org.apache.spark.ml.util._
import org.apache.spark.sql.types.{DataType, StringType}

class UnicodeNormalizer (override val uid: String)
  extends UnaryTransformer[String, String, UnicodeNormalizer] {

  def this() = this(Identifiable.randomUID("unicode_normalizer"))

  private val forms = Map(
    "NFC" -> Normalizer.Form.NFC, "NFD" -> Normalizer.Form.NFD,
    "NFKC" -> Normalizer.Form.NFKC, "NFKD" -> Normalizer.Form.NFKD
  )

  val form: Param[String] = new Param(this, "form", "unicode form (one of NFC, NFD, NFKC, NFKD)",
    ParamValidators.inArray(forms.keys.toArray))

  def setN(value: String): this.type = set(form, value)

  def getForm: String = $(form)

  setDefault(form -> "NFKD")

  override protected def createTransformFunc: String => String = {
    val normalizerForm = forms($(form))
    (s: String) => Normalizer.normalize(s, normalizerForm)
  }

  override protected def validateInputType(inputType: DataType): Unit = {
    require(inputType == StringType, s"Input type must be string type but got $inputType.")
  }

  override protected def outputDataType: DataType = StringType
}

相应的构建定义(调整 Spark 和 Scala 版本以匹配您的 Spark 部署):

name := "unicode-normalization"

version := "1.0"

crossScalaVersions := Seq("2.11.12", "2.12.8")

organization := "net.zero323"

val sparkVersion = "2.4.0"

libraryDependencies ++= Seq(
  "org.apache.spark" %% "spark-core" % sparkVersion,
  "org.apache.spark" %% "spark-sql" % sparkVersion,
  "org.apache.spark" %% "spark-mllib" % sparkVersion
)

在 Python 方面，您需要一个类似于此的包装器。

from pyspark.ml.param.shared import *
# from pyspark.ml.util import keyword_only  # in Spark < 2.0
from pyspark import keyword_only 
from pyspark.ml.wrapper import JavaTransformer

class UnicodeNormalizer(JavaTransformer, HasInputCol, HasOutputCol):

    @keyword_only
    def __init__(self, form="NFKD", inputCol=None, outputCol=None):
        super(UnicodeNormalizer, self).__init__()
        self._java_obj = self._new_java_obj(
            "net.zero323.spark.ml.feature.UnicodeNormalizer", self.uid)
        self.form = Param(self, "form",
            "unicode form (one of NFC, NFD, NFKC, NFKD)")
        # kwargs = self.__init__._input_kwargs  # in Spark < 2.0
        kwargs = self._input_kwargs
        self.setParams(**kwargs)

    @keyword_only
    def setParams(self, form="NFKD", inputCol=None, outputCol=None):
        # kwargs = self.setParams._input_kwargs  # in Spark < 2.0
        kwargs = self._input_kwargs
        return self._set(**kwargs)

    def setForm(self, value):
        return self._set(form=value)

    def getForm(self):
        return self.getOrDefault(self.form)

构建 Scala 包:

sbt +package

在启动 shell 或提交时包含它。例如，对于使用 Scala 2.11 构建的 Spark:

bin/pyspark --jars path-to/target/scala-2.11/unicode-normalization_2.11-1.0.jar \
 --driver-class-path path-to/target/scala-2.11/unicode-normalization_2.11-1.0.jar

你应该准备好了。剩下的就是一些正则表达式的魔法:

from pyspark.sql.functions import regexp_replace

normalizer = UnicodeNormalizer(form="NFKD",
    inputCol="text", outputCol="text_normalized")

df = sc.parallelize([
    (1, "Maracaibó"), (2, "New York"),
    (3, "   São Paulo   "), (4, "~Madrid")
]).toDF(["id", "text"])

(normalizer
    .transform(df)
    .select(regexp_replace("text_normalized", "\p{M}", ""))
    .show())

## +--------------------------------------+
## |regexp_replace(text_normalized,\p{M},)|
## +--------------------------------------+
## |                             Maracaibo|
## |                              New York|
## |                          Sao Paulo   |
## |                               ~Madrid|
## +--------------------------------------+

请注意，这遵循与内置文本转换器相同的约定，并且不安全。您可以通过检查 createTransformFunc 中的 null 轻松纠正该问题。

关于python - 在 PySpark 中使用 Apache Spark 数据帧删除重音的最佳方法是什么？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/38359534/

文章推荐： c - 在 Eclipse for C 中包含外部 header

文章推荐： c - 如何解析两个同名的结构体？

文章推荐： python - 如何跟踪玩家的排名？

PostgreSQL 重音 + 不区分大小写的搜索
我正在寻找一种方法来支持不区分大小写 + 重音不区分搜索的良好性能。到目前为止，我们在使用 MSSql 服务器时没有遇到任何问题，在 Oracle 上我们必须使用 OracleText，而现在我们在
php - 重音 "e"即使在元标记之后也显示为问号
这个问题已经有答案了: Trouble with UTF-8 characters; what I see is not what I stored (5 个回答) 已关闭 5 年前。我刚刚将一个我
linux - 使用反引号/重音/波形符作为修饰键
我正在寻找一种在 Linux 中使用反引号 (`)/波形符 (~) 键和其他一些键创建键盘快捷键的方法。在理想情况下: 按下波形符没有任何作用按下波形符的同时按另一个键会触发(可自定义的)快捷方式
php preg_grep 和元音变音/重音
我有一个由术语组成的数组，其中一些包含重音字符。我像这样做一个 preg grep $data= array('Napoléon','Café'); $result = preg_grep('~' .
.net - DataGridView 过滤器忽略单元格、单词上的变音符号(重音)
我使用 TextBox 在 DataGridView 中进行过滤 image .这是完美的工作。表格的单元格包含 1250 个拉丁字符。我想搜索忽略单元格中单词的重音。例子。如果是文本框 "knjaz
vim - .vimrc 中的键映射(重音)和编码问题
我在 Vim 中遇到一个奇怪的映射问题。我使用的是 Azerty 键盘。在我的 .vimrc 中，我有以下命令可以在段落之间快速移动。 nnoremap _ { vnoremap _ { nnore
javascript - nodejs 中的 Utf8 重音
我尝试读取一个utf8编码的vcf文件，结果是: { "name": "=4A=61=76=69=65=72=20=4C=75=6A=C3=A1=6E", "tel":
mysql - 奇怪的 MYSQL 反引号(重音)
我的数据库中有两个表，info 和 comment，它们的结构如下: info (id(int(10)), name(varchar(80)), ...19 other columns.., phon
linux - Linux 中的 QtWebkit 重音
我使用 QtWebkit 制作了一个应用程序。在同一个 html 页面中，在 Windows 上使用重音符号(西类牙语)时可以正常工作，但在 Linux (Ubuntu) 上则不起作用。我不明白为什
php - 比较两个字符串并忽略(但不替换)重音。 PHP
我有(例如)两个字符串: $a = "joao"; $b = "joão"; if ( strtoupper($a) == strtoupper($b)) { echo $b; } 我希望它是
ruby - 将法语(重音)字符放入 Ruby 文件中
这个问题在这里已经有了答案: 关闭 10 年前。 Possible Duplicate: invalid multibyte char (US-ASCII) with Rails and Ruby
php - 重写 'pretty URLs' 时如何处理变音符号(重音)
我重写 URL 以包含用户生成的旅游博客的标题。我这样做是为了 URL 的可读性和 SEO 目的。 http://www.example.com/gallery/280-Gorges_du_Tod
c++ - 如何使用 ncurses 获取 UTF-8 重音
我最近安装了新的 Windows 10 build 14393，我想使用新的 linux 子系统。所以我决定学习 ncurses，但我找不到如何从 getch 中获取带有重音符的字符的 UTF-8 代

太空狗

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - 在 PySpark 中使用 Apache Spark 数据帧删除重音的最佳方法是什么？