- html - 出于某种原因,IE8 对我的 Sass 文件中继承的 html5 CSS 不友好?
- JMeter 在响应断言中使用 span 标签的问题
- html - 在 :hover and :active? 上具有不同效果的 CSS 动画
- html - 相对于居中的 html 内容固定的 CSS 重复背景?
Apache Spark StringIndexerModel 在对某一特定列进行转换后返回空数据集。我正在使用成人数据集:http://mlr.cs.umass.edu/ml/datasets/Adult
第1步:创建StringIndexerModel并保存到本地
StringIndexerModel model = new StringIndexer().setInputCol(column).setOutputCol("label").setHandleInvalid("skip").setStringOrderType("alphabetAsc").fit(originalDataset);
model.write().save(filelocation);
第 2 步:读取索引器模型并转换新数据集
StringIndexerModel model = StringIndexerModel.read().load(filelocation);
newDataset = model.transform(newDataset).drop(column).withColumnRenamed("label", column);
新数据集:
+---+------------+------------+----------+-------------+------+--------------+-------------------+--------------+----------------+-----+--------------+----+-----------------+
|age|capital gain|capital loss|education |education num|fnlgwt|hours per week|marital status |native country|occupation |race |relationship |sex |workclass |
+---+------------+------------+----------+-------------+------+--------------+-------------------+--------------+----------------+-----+--------------+----+-----------------+
|39 |2174 |0 | Bachelors|13 |77516 |40 | Never-married | United-States| Adm-clerical |White| Not-in-family|Male| State-gov |
|50 |0 |0 | Bachelors|13 |83311 |13 | Married-civ-spouse| United-States| Exec-managerial|White| Husband |Male| Self-emp-not-inc|
+---+------------+------------+----------+-------------+------+--------------+-------------------+--------------+----------------+-----+--------------+----+-----------------+
正确输出:
Column: education | File Location: localFolder/stringIndex/education
Labels: [ 10th, 11th, 12th, 1st-4th, 5th-6th, 7th-8th, 9th, Assoc-acdm, Assoc-voc, Bachelors, Doctorate, HS-grad, Masters, Preschool, Prof-school, Some-college]
+---+------------+------------+-------------+------+--------------+-------------------+--------------+----------------+-----+--------------+----+-----------------+---------+
|age|capital gain|capital loss|education num|fnlgwt|hours per week|marital status |native country|occupation |race |relationship |sex |workclass |education|
+---+------------+------------+-------------+------+--------------+-------------------+--------------+----------------+-----+--------------+----+-----------------+---------+
|39 |2174 |0 |13 |77516 |40 | Never-married | United-States| Adm-clerical |White| Not-in-family|Male| State-gov |9.0 |
|50 |0 |0 |13 |83311 |13 | Married-civ-spouse| United-States| Exec-managerial|White| Husband |Male| Self-emp-not-inc|9.0 |
+---+------------+------------+-------------+------+--------------+-------------------+--------------+----------------+-----+--------------+----+-----------------+---------+
Column: marital status | File Location: localFolder/stringIndex/marital status
Labels: [ Divorced, Married-AF-spouse, Married-civ-spouse, Married-spouse-absent, Never-married, Separated, Widowed]
+---+------------+------------+-------------+------+--------------+--------------+----------------+-----+--------------+----+-----------------+---------+--------------+
|age|capital gain|capital loss|education num|fnlgwt|hours per week|native country|occupation |race |relationship |sex |workclass |education|marital status|
+---+------------+------------+-------------+------+--------------+--------------+----------------+-----+--------------+----+-----------------+---------+--------------+
|39 |2174 |0 |13 |77516 |40 | United-States| Adm-clerical |White| Not-in-family|Male| State-gov |9.0 |4.0 |
|50 |0 |0 |13 |83311 |13 | United-States| Exec-managerial|White| Husband |Male| Self-emp-not-inc|9.0 |2.0 |
+---+------------+------------+-------------+------+--------------+--------------+----------------+-----+--------------+----+-----------------+---------+--------------+
Column: native country | File Location: localFolder/stringIndex/native country
Labels: [ ?, Cambodia, Canada, China, Columbia, Cuba, Dominican-Republic, Ecuador, El-Salvador, England, France, Germany, Greece, Guatemala, Haiti, Holand-Netherlands, Honduras, Hong, Hungary, India, Iran, Ireland, Italy, Jamaica, Japan, Laos, Mexico, Nicaragua, Outlying-US(Guam-USVI-etc), Peru, Philippines, Poland, Portugal, Puerto-Rico, Scotland, South, Taiwan, Thailand, Trinadad&Tobago, United-States, Vietnam, Yugoslavia]
+---+------------+------------+-------------+------+--------------+----------------+-----+--------------+----+-----------------+---------+--------------+--------------+
|age|capital gain|capital loss|education num|fnlgwt|hours per week|occupation |race |relationship |sex |workclass |education|marital status|native country|
+---+------------+------------+-------------+------+--------------+----------------+-----+--------------+----+-----------------+---------+--------------+--------------+
|39 |2174 |0 |13 |77516 |40 | Adm-clerical |White| Not-in-family|Male| State-gov |9.0 |4.0 |39.0 |
|50 |0 |0 |13 |83311 |13 | Exec-managerial|White| Husband |Male| Self-emp-not-inc|9.0 |2.0 |39.0 |
+---+------------+------------+-------------+------+--------------+----------------+-----+--------------+----+-----------------+---------+--------------+--------------+
Column: occupation | File Location: localFolder/stringIndex/occupation
Labels: [ ?, Adm-clerical, Armed-Forces, Craft-repair, Exec-managerial, Farming-fishing, Handlers-cleaners, Machine-op-inspct, Other-service, Priv-house-serv, Prof-specialty, Protective-serv, Sales, Tech-support, Transport-moving]
+---+------------+------------+-------------+------+--------------+-----+--------------+----+-----------------+---------+--------------+--------------+----------+
|age|capital gain|capital loss|education num|fnlgwt|hours per week|race |relationship |sex |workclass |education|marital status|native country|occupation|
+---+------------+------------+-------------+------+--------------+-----+--------------+----+-----------------+---------+--------------+--------------+----------+
|39 |2174 |0 |13 |77516 |40 |White| Not-in-family|Male| State-gov |9.0 |4.0 |39.0 |1.0 |
|50 |0 |0 |13 |83311 |13 |White| Husband |Male| Self-emp-not-inc|9.0 |2.0 |39.0 |4.0 |
+---+------------+------------+-------------+------+--------------+-----+--------------+----+-----------------+---------+--------------+--------------+----------+
输出错误:除此之外所有其他模型都工作正常
Column: race | File Location: localFolder/stringIndex/race
Labels: [ Amer-Indian-Eskimo, Asian-Pac-Islander, Black, Other, White]
+---+------------+------------+-------------+------+--------------+------------+---+---------+---------+--------------+--------------+----------+----+
|age|capital gain|capital loss|education num|fnlgwt|hours per week|relationship|sex|workclass|education|marital status|native country|occupation|race|
+---+------------+------------+-------------+------+--------------+------------+---+---------+---------+--------------+--------------+----------+----+
+---+------------+------------+-------------+------+--------------+------------+---+---------+---------+--------------+--------------+----------+----+
如果您能帮助解决此问题,我将不胜感激。谢谢!
最佳答案
事实证明,新数据集的数据不正确。值之前应有空格。
添加空格'White'
让我得到了正确的输出。
关于java - Spark StringIndexer 返回空数据集,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/59518208/
我已经为使用 JGroups 编写了简单的测试。有两个像这样的简单应用程序 import org.jgroups.*; import org.jgroups.conf.ConfiguratorFact
我有一个通过 ajax 检索的 json 编码数据集。我尝试检索的一些数据点将返回 null 或空。 但是,我不希望将那些 null 或空值显示给最终用户,或传递给其他函数。 我现在正在做的是检查
这个问题在这里已经有了答案: 关闭 11 年前。 Possible Duplicate: Why does one often see “null != variable” instead of “
嗨在我们公司,他们遵循与空值进行比较的严格规则。当我编码 if(variable!=null) 在代码审查中,我收到了对此的评论,将其更改为 if(null!=variable)。上面的代码对性能有影
我正在尝试使用 native Cordova QR 扫描仪插件编译项目,但是我不断收到此错误。据我了解,这是代码编写方式的问题,它向构造函数发送了错误的值,或者根本就没有找到构造函数。那么我该如何解决
我在装有 Java 1.8 的 Windows 10 上使用 Apache Nutch 1.14。我已按照 https://wiki.apache.org/nutch/NutchTutorial 中提
这个问题已经有答案了: 已关闭11 年前。 Possible Duplicate: what is “=null” and “ IS NULL” Is there any difference bet
Three-EyedRaven 内网渗透初期,我们都希望可以豪无遗漏的尽最大可能打开目标内网攻击面,故,设计该工具的初衷是解决某些工具内网探测速率慢、运行卡死、服务爆破误报率高以及socks流
我想在Scala中像在Java中那样做: public void recv(String from) { recv(from, null); } public void recv(String
我正在尝试从一组图像补丁中创建一个密码本。我已将图像(Caltech 101)分成20 X 20图像块。我想为每个补丁创建一个SIFT描述符。但是对于某些图像补丁,它不返回任何描述符/关键点。我尝试使
我在验证器类中自动连接的两个服务有问题。这些服务工作正常,因为在我的 Controller 中是自动连接的。我有一个 applicationContext.xml 文件和 MyApp-servlet.
已关闭。此问题不符合Stack Overflow guidelines 。目前不接受答案。 已关闭10 年前。 问题必须表现出对要解决的问题的最低程度的了解。告诉我们您尝试过做什么,为什么不起作用,以
大家好,我正在对数据库进行正常的选择,但是 mysql_num_rowsis 为空,我不知道为什么,我有 7 行选择。 如果您发现问题,请告诉我。 真的谢谢。 代码如下: function get_b
我想以以下格式创建一个字符串:id[]=%@&stringdata[]=%@&id[]=%@&stringdata[]=%@&id[]=%@&stringdata[]=%@&等,在for循环中,我得到
我正在尝试使用以下代码将URL转换为字符串: NSURL *urlOfOpenedFile = _service.myURLRequest.URL; NSString *fileThatWasOpen
我正在尝试将NSNumber传递到正在工作的UInt32中。然后,我试图将UInt32填充到NSData对象中。但是,这在这里变得有些时髦... 当我尝试将NSData对象中的内容写成它返回的字符串(
我正在进行身份验证并收到空 cookie。我想存储这个 cookie,但服务器没有返回给我 cookie。但响应代码是 200 ok。 httpConn.setRequestProperty(
我认为 Button bTutorial1 = (Button) findViewById(R.layout.tutorial1); bTutorial1.setOnClickListener
我的 Controller 中有这样的东西: model.attribute("hiringManagerMap",hiringManagerMap); 我正在访问此 hiringManagerMap
我想知道如何以正确的方式清空列表。在 div 中有一个列表然后清空 div 或列表更好吗? 我知道这是一个蹩脚的问题,但请帮助我理解这个 empty() 函数:) 案例)如果我运行这个脚本会发生什么:
我是一名优秀的程序员,十分优秀!