python - 具有两个分类变量的 Matplotlib 点图-6ren

python - 具有两个分类变量的 Matplotlib 点图

转载作者：行者123 更新时间：2023-12-01 00:54:08

我想生成一种特定类型的可视化，包含一个相当简单的 dot plot但有一点不同:两个轴都是分类变量(即序数或非数值)。这不但没有让事情变得更容易，反而让事情变得更加复杂。

为了说明这个问题，我将使用一个小型示例数据集，该数据集是对 seaborn.load_dataset("tips") 的修改，并定义如下:

import pandas
from six import StringIO
df = """total_bill |  tip  |    sex | smoker | day |   time | size
             16.99 | 1.01  |   Male |     No | Mon | Dinner |    2
             10.34 | 1.66  |   Male |     No | Sun | Dinner |    3
             21.01 | 3.50  |   Male |     No | Sun | Dinner |    3
             23.68 | 3.31  |   Male |     No | Sun | Dinner |    2
             24.59 | 3.61  | Female |     No | Sun | Dinner |    4
             25.29 | 4.71  | Female |     No | Mon | Lunch  |    4
              8.77 | 2.00  | Female |     No | Tue | Lunch  |    2
             26.88 | 3.12  |   Male |     No | Wed | Lunch  |    4
             15.04 | 3.96  |   Male |     No | Sat | Lunch  |    2
             14.78 | 3.23  |   Male |     No | Sun | Lunch  |    2"""
df = pandas.read_csv(StringIO(df.replace(' ','')), sep="|", header=0)

生成图表的第一种方法是尝试调用 seaborn，如下所示:

import seaborn
axes = seaborn.pointplot(x="time", y="sex", data=df)

此操作失败:

ValueError: Neither the `x` nor `y` variable appears to be numeric.

等效的 seaborn.stripplot 和 seaborn.swarmplot 调用也是如此。但是，如果其中一个变量是分类变量而另一个变量是数值变量，则它确实有效。确实 seaborn.pointplot(x="total_bill", y="sex", data=df) 有效，但不是我想要的。

我还尝试了这样的散点图:

axes = seaborn.scatterplot(x="time", y="sex", size="day", data=df,
                           x_jitter=True, y_jitter=True)

这会产生以下图表，该图表不包含任何抖动，并且所有点都重叠，因此毫无用处:

你知道有什么优雅的方法或库可以解决我的问题吗？

我开始自己写一些东西，我将在下面包含它，但这种实现不是最理想的，并且受到可以在同一点重叠的点的数量的限制(目前，如果超过 4 个点重叠，它就会失败)。

# Modules #
import seaborn, pandas, matplotlib
from six import StringIO

################################################################################
def amount_to_offets(amount):
    """A function that takes an amount of overlapping points (e.g. 3)
    and returns a list of offsets (jittered) coordinates for each of the
    points.

    It follows the logic that two points are displayed side by side:

    2 ->  * *

    Three points are organized in a triangle

    3 ->   *
          * *

    Four points are sorted into a square, and so on.

    4 ->  * *
          * *
    """
    assert isinstance(amount, int)
    solutions = {
        1: [( 0.0,  0.0)],
        2: [(-0.5,  0.0), ( 0.5,  0.0)],
        3: [(-0.5, -0.5), ( 0.0,  0.5), ( 0.5, -0.5)],
        4: [(-0.5, -0.5), ( 0.5,  0.5), ( 0.5, -0.5), (-0.5,  0.5)],
    }
    return solutions[amount]

################################################################################
class JitterDotplot(object):

    def __init__(self, data, x_col='time', y_col='sex', z_col='tip'):
        self.data = data
        self.x_col = x_col
        self.y_col = y_col
        self.z_col = z_col

    def plot(self, **kwargs):
        # Load data #
        self.df = self.data.copy()

        # Assign numerical values to the categorical data #
        # So that ['Dinner', 'Lunch'] becomes [0, 1] etc. #
        self.x_values = self.df[self.x_col].unique()
        self.y_values = self.df[self.y_col].unique()
        self.x_mapping = dict(zip(self.x_values, range(len(self.x_values))))
        self.y_mapping = dict(zip(self.y_values, range(len(self.y_values))))
        self.df = self.df.replace({self.x_col: self.x_mapping, self.y_col: self.y_mapping})

        # Offset points that are overlapping in the same location #
        # So that (2.0, 3.0) becomes (2.05, 2.95) for instance #
        cols = [self.x_col, self.y_col]
        scaling_factor = 0.05
        for values, df_view in self.df.groupby(cols):
            offsets = amount_to_offets(len(df_view))
            offsets = pandas.DataFrame(offsets, index=df_view.index, columns=cols)
            offsets *= scaling_factor
            self.df.loc[offsets.index, cols] += offsets

        # Plot a standard scatter plot #
        g = seaborn.scatterplot(x=self.x_col, y=self.y_col, size=self.z_col, data=self.df, **kwargs)

        # Force integer ticks on the x and y axes #
        locator = matplotlib.ticker.MaxNLocator(integer=True)
        g.xaxis.set_major_locator(locator)
        g.yaxis.set_major_locator(locator)
        g.grid(False)

        # Expand the axis limits for x and y #
        margin = 0.4
        xmin, xmax, ymin, ymax = g.get_xlim() + g.get_ylim()
        g.set_xlim(xmin-margin, xmax+margin)
        g.set_ylim(ymin-margin, ymax+margin)

        # Replace ticks with the original categorical names #
        g.set_xticklabels([''] + list(self.x_mapping.keys()))
        g.set_yticklabels([''] + list(self.y_mapping.keys()))

        # Return for display in notebooks for instance #
        return g

################################################################################
# Graph #
graph = JitterDotplot(data=df)
axes  = graph.plot()
axes.figure.savefig('jitter_dotplot.png')

最佳答案

您可以首先将时间和性别转换为分类类型并稍微调整一下:

df.sex = pd.Categorical(df.sex)
df.time = pd.Categorical(df.time)

axes = sns.scatterplot(x=df.time.cat.codes+np.random.uniform(-0.1,0.1, len(df)), 
                       y=df.sex.cat.codes+np.random.uniform(-0.1,0.1, len(df)),
                       size=df.tip)

输出:

有了这个想法，您可以将上述代码中的偏移量(np.random)修改为相应的距离。例如:

# grouping
groups = df.groupby(['time', 'sex'])

# compute the number of samples per group
num_samples = groups.tip.transform('size')

# enumerate the samples within a group
sample_ranks = df.groupby(['time']).cumcount() * (2*np.pi) / num_samples

# compute the offset
x_offsets = np.where(num_samples.eq(1), 0, np.cos(df.sample_rank) * 0.03)
y_offsets = np.where(num_samples.eq(1), 0, np.sin(df.sample_rank) * 0.03)

# plot
axes = sns.scatterplot(x=df.time.cat.codes + x_offsets, 
                       y=df.sex.cat.codes + y_offsets,
                       size=df.tip)

输出:

关于python - 具有两个分类变量的 Matplotlib 点图，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/56347325/

文章推荐： javascript - 在循环内等待 API 的结果

文章推荐： jquery - 编写移动网站的最佳方法？

文章推荐： sbt 创建多个 scala 源目录

matplotlib - matplotlib 中对数极坐标图轴标签的定位
我无法在此图中定位轴标签。我喜欢放置顶部标签，使管道与网格对齐，并放置左右标签，以便它们不接触绘图。我试过了 ax.tick_params(axis='both', which='both'
matplotlib - matplotlib 中的条形图应该如何设置宽度？
我使用的是 python 2，下面的代码只是使用了一些示例数据，我的实际数据可能有不同的长度，并且可能不是很细。 import numpy as np import datetime i
matplotlib - Matplotlib 中的线段
给定坐标 [1,5,7,3,5,10,3,6,8]为 matplotlib.pyplot ，如何突出显示或着色线条的不同部分。例如，列表中的坐标 1-3 ( [1,5,7,3] ) 表示属性 a .我
matplotlib - Matplotlib 3D绘图中的较深背景
我正在matplotlib中绘制以下图像。我的问题是，图像看起来像这样，但是，我想使背景变暗，因为当我打印该图像时，灰度部分不会出现在打印物中。有人可以告诉我API进行此更改吗？我使用简单的API
matplotlib - matplotlib，逐步动画
这是关于matplotlib的一个非常基本的问题，但是我不知道该怎么做: 我想绘制多个图形，并使用绘制窗口中的箭头从一个移到另一个。目前，我只知道如何创建多个图并将其绘制在不同的窗口中，如下所示:
matplotlib - matplotlib 补丁绘图中的工件
在 matplotlib 中绘制小块对象时，由于显示分辨率而引入了伪影。使用抗锯齿并不能解决问题。这个问题有解决方案吗？ import matplotlib.pyplot as plt impo
matplotlib - matplotlib 中的未填充条形图
对于直方图，有一个简单的内置选项 histtype='step' .如何制作相同风格的条形图？最佳答案 [阅读评论后添加答案] 将可选关键字设置为 fill=False对于条形图: import m
matplotlib - matplotlib 子图中的图例位置
我正在尝试在 (6X3) 网格上创建子图。我对图例的位置有疑问。图例对所有子图都是通用的。 lgend 现在与 y 轴标签重叠我尝试删除 constrained_layout=True 选项。但这在
matplotlib - matplotlib 中的点和线工具提示？
我有一个带有一些线段( LineCollection )和一些点的图表。这些线和点有一些与它们相关的值，但没有绘制出来。我希望能够添加鼠标悬停工具提示或其他方法来轻松找到点和线的关联值。这对于点或线段
matplotlib - Matplotlib 图图例中的制表符对齐
我想创建一个带有对齐不同曲线文本的图例的图。这是一个最小的工作示例: import matplotlib.pyplot as plt import numpy as np x=np.linspace(
matplotlib - Matplotlib:图例中的水平线长
可以说我正在用matplotlib绘制一条线并添加一个图例。在图例中，其显示为------ Label。当绘制较小的图形尺寸以进行打印时，我发现该行的默认水平长度太长。是否存在将------ La
matplotlib - matplotlib 图形中的常见起源
我正在使用 matplotlib 构建一个 3D 散点图，但无法使生成的图形具有所有 3 个轴的共同原点。我怎样才能做到这一点？我的代码(到目前为止)，我还没有为轴规范实现任何定义，因为我对 Pyt
matplotlib - matplotlib 中是否存在用于在子图中定义子图网格的工具？
我有一个我想使用的绘图布局，其中 9 个不同的数据簇被布置在一个方形网格上。网格中的每个框都包含 3 个并排布置的箱线图。我最初的想法是这将适合 3x3 子图布局，每个单独的子图本身被划分为 3x1
matplotlib - Matplotlib，如何在数据坐标之外的图形外部编写注释？
我的图形从y=-1变为y=10 我想在任意位置写一小段文字，例如x=2000，y=5: ax.annotate('MgII', xy=(2000.0, 5.0), xycoords='data')
matplotlib - Matplotlib-在LateX表达式中使用变量
我想使用LateX格式来构建一个表达式，其中出现一些数字，但这些数字是用LateX表达式中的变量表示的。实际的目标是在axes.annotate()方法中使用它，但是为了讨论起见，这里是一个原理代码
matplotlib - Matplotlib 中的叠加轮廓图
我需要比较两组的二维分布。当我使用 matplotlib.pyplot.contourf并覆盖图，每个等高线图的背景颜色填充整个图空间。有没有办法让每个等高线图的最低等高线级别透明，以便更容易看到每
matplotlib - matplotlib —以交互方式选择点或位置？
在R中，有一个locator函数，类似于Matlab的ginput，您可以用鼠标单击图形并选择任何x，y坐标。此外，还有一个名为identify(x,y)的函数，如果您给它绘制了一组绘制的点x，y，然
matplotlib - matplotlib:生成矢量图
我想用matplotlib生成矢量图。我尽力了-但输出是光栅图像。这是我使用的： import matplotlib matplotlib.use('Agg') import matplotlib.p
matplotlib - matplotlib 中的小散点图标记始终为黑色
我正在尝试使用 matplotlib 制作具有非常小的灰点的散点图。由于点密度的原因，点需要很小。问题是 scatter() 函数的标记似乎既有线条又有填充。当标记很小时，只有线条可见，而看不到填充，
matplotlib - matplotlib 中的垂直线和水平线
我不太明白为什么我无法在指定的限制内创建水平和垂直线。我想用这个框绑定(bind)数据。然而，双方似乎并没有遵守我的指示。为什么是这样？ # CREATING A BOUNDING BOX # BOT

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - 具有两个分类变量的 Matplotlib 点图