python - 创建可用值的分布 - Python

转载作者：太空宇宙更新时间：2023-11-04 09:32:06

我希望创建一个分布来显示员工可以上类的时间。类似于此图，可在此链接中找到 staff distribution .

为实现这一点，我创建了 staff_availability_df，其中包含要从中挑选的员工数量，可在 ['Person'] 列中找到。他们可以工作的 min - max 小时数，他们得到的报酬是多少都有这样的标签。 The available times they can work are separated into hours ['Availability_Hr']，表示他们可以工作的时间，以小时表示。所以第一人称是'8-18'，也就是8:00:00am - 18:00:00pm。 ['Availability_15min_Seg'] 本质上是相同的，但小时分为 4 个部分。所以第一人称是 '1-41'，又是 8:00:00am - 18:00:00pm。

注意:标准类次在 8:00:00am - 3:30:00am 之间运行，因此大约 20 小时。

staff_requirements_df 显示整个类次的时间 和我需要的所需人员。

import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.dates as dates

#This is the employee availability:
staff_availability = pd.DataFrame({
    'Person' : ['C1','C2','C3','C4','C5','C6','C7','C8','C9','C10','C11'],                 
    'MinHours' : [5,5,5,5,5,5,5,5,5,5,5],    
    'MaxHours' : [10,10,10,10,10,10,10,10,10,10,10],                 
    'HourlyWage' : [26,26,26,26,26,26,26,26,26,26,26],  
    'Availability_Hr' : ['8-18','8-18','8-18','9-18','9-18','9-18','12-1','12-1','17-3','17-3','17-3'],                              
    'Availability_15min_Seg' : ['1-41','1-41','1-41','5-41','5-41','5-41','17-69','17-69','37-79','37-79','37-79'],                              
    }) 

#These are the staffing requirements:
staffing_requirements = pd.DataFrame({
    'Time' : ['0/1/1900 8:00:00','0/1/1900 9:59:00','0/1/1900 10:00:00','0/1/1900 12:29:00','0/1/1900 12:30:00','0/1/1900 13:00:00','0/1/1900 13:02:00','0/1/1900 13:15:00','0/1/1900 13:20:00','0/1/1900 18:10:00','0/1/1900 18:15:00','0/1/1900 18:20:00','0/1/1900 18:25:00','0/1/1900 18:45:00','0/1/1900 18:50:00','0/1/1900 19:05:00','0/1/1900 19:07:00','0/1/1900 21:57:00','0/1/1900 22:00:00','0/1/1900 22:30:00','0/1/1900 22:35:00','1/1/1900 3:00:00','1/1/1900 3:05:00','1/1/1900 3:20:00','1/1/1900 3:25:00'],                 
    'People' : [1,1,2,2,3,3,2,2,3,3,4,4,3,3,2,2,3,3,4,4,3,3,2,2,1],                      
     })

我使用以下函数导出了发生在 8:00:00am - 3:30:00am 之间的 15 分钟片段中的人员配备要求。每 15 分钟分配给 string 'T'。所以 T1 = 8:00:00am 和 T79 = 3:00:00am

staffing_requirements['Time'] = ['/'.join([str(int(x.split('/')[0])+1)] + x.split('/')[1:]) for x in staffing_requirements['Time']]
staffing_requirements['Time'] = pd.to_datetime(staffing_requirements['Time'], format='%d/%m/%Y %H:%M:%S')
formatter = dates.DateFormatter('%Y-%m-%d %H:%M:%S') 

staffing_requirements = staffing_requirements.groupby(pd.Grouper(freq='15T',key='Time'))['People'].max().ffill()
staffing_requirements = staffing_requirements.reset_index(level=['Time'])

staffing_requirements.insert(2, 'T', range(1, 1 + len(staffing_requirements)))
staffing_requirements['T'] = 'T' + staffing_requirements['T'].astype(str)

st_req = staffing_requirements['People'].tolist()

[1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 2.0, 2.0, 2.0, 2.0, 2.0, 2.0, 2.0, 2.0, 2.0, 2.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 4.0, 4.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 3.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 3.0, 2.0]

我希望使用这些函数来创建一个线性规划矩阵，它返回每个员工可以工作的时间的分布。但我希望使用 15 分钟的片段以及几个小时。例如注意:此导出将延长至 3:30am。所以它将包含 79 个段。

注意:要清楚。我希望返回分发时间表，以便将来使用。不仅仅是一个数字。

有几个工作人员可用example 1 example 2使用混合整数线性规划 的方法，但它们使用闭源软件。我希望将其翻译成 Python。

最佳答案

这对于整数规划来说确实是一项伟大的工作；您可以使用 pulp，您首先需要通过命令行安装它，例如pip install pulp

数据操纵为成功做好准备

然后，首先确保您的 DataFrames 处于最佳状态，以便我们可以解决问题:

# Since timeslots for staffing start counting at 1, also make the
# DataFrame index start counting at 1
staffing_requirements.index = range(1, len(staffing_requirements) + 1) 
print(staffing_requirements.tail())

staff_availability.set_index('Person')

staff_costs = staff_availability.set_index('Person')[['MinHours', 'MaxHours', 'HourlyWage']]
availability = staff_availability.set_index('Person')[['Availability_15min_Seg']]
availability[['first_15min', 'last_15min']] =  availability['Availability_15min_Seg'].str.split('-', expand=True).astype(int)

availability_per_member =  [pd.DataFrame(1, columns=[idx], index=range(row['first_15min'], row['last_15min']+1))
 for idx, row in availability.iterrows()]

availability_per_member = pd.concat(availability_per_member, axis='columns').fillna(0).astype(int).stack()
availability_per_member.index.names = ['Timeslot', 'Person']
availability_per_member = (availability_per_member.to_frame()
                            .join(staff_costs[['HourlyWage']])
                            .rename(columns={0: 'Available'}))

其中 availability_per_member 现在是一个 MultiIndex DataFrame，每个人每个时间段一行，指示他/她的可用性和工资:

#                 Available  HourlyWage
#Timeslot Person                       
#1        C1              1          26
#         C2              1          26
#         C3              1          26
#         C4              0          26
#         C5              0          26

此外，我们稍微改变了先决条件，使问题实际上是可以解决的；请参阅附录了解为什么这是必要的

import numpy as np
np.random.seed(42)
staffing_requirements['People'] = np.random.randint(1, 4, size=len(staffing_requirements))
staff_costs['MinHours'] = 3

用`pulp`解决整数规划问题

现在，我们可以让 pulp 开始工作了:以最小化成本为目标设置问题，并逐一添加您提到的约束，请参阅注释代码。 staffed 现在是一个 pulp-dictionary，其中包含一个人是否在某个时间段(0 或 1)有工作人员

import pulp
prob = pulp.LpProblem('CreateStaffing', pulp.LpMinimize) # Minimize costs

timeslots = staffing_requirements.index
persons = availability_per_member.index.levels[1]

# A member is either staffed or is not at a certain timeslot
staffed = pulp.LpVariable.dicts("staffed",
                                   ((timeslot, staffmember) for timeslot, staffmember 
                                    in availability_per_member.index),
                                     lowBound=0,
                                     cat='Binary')

# Objective = cost (= sum of hourly wages)                              
prob += pulp.lpSum(
    [staffed[timeslot, staffmember] * availability_per_member.loc[(timeslot, staffmember), 'HourlyWage'] 
    for timeslot, staffmember in availability_per_member.index]
)

# Staff the right number of people
for timeslot in timeslots:
    prob += (sum([staffed[(timeslot, person)] for person in persons]) 
    == staffing_requirements.loc[timeslot, 'People'])


# Do not staff unavailable persons
for timeslot in timeslots:
    for person in persons:
        if availability_per_member.loc[(timeslot, person), 'Available'] == 0:
            prob += staffed[timeslot, person] == 0

# Do not underemploy people
for person in persons:
    prob += (sum([staffed[(timeslot, person)] for timeslot in timeslots])
    >= staff_costs.loc[person, 'MinHours']*4) # timeslot is 15 minutes => 4 timeslots = hour

# Do not overemploy people
for person in persons:
    prob += (sum([staffed[(timeslot, person)] for timeslot in timeslots])
    <= staff_costs.loc[person, 'MaxHours']*4) # timeslot is 15 minutes => 4 timeslots = hour

然后，就是让 pulp 解决这个问题了:

prob.solve()
print(pulp.LpStatus[prob.status])

output = []
for timeslot, staffmember in staffed:
    var_output = {
        'Timeslot': timeslot,
        'Staffmember': staffmember,
        'Staffed': staffed[(timeslot, staffmember)].varValue,
    }
    output.append(var_output)
output_df = pd.DataFrame.from_records(output)#.sort_values(['timeslot', 'staffmember'])
output_df.set_index(['Timeslot', 'Staffmember'], inplace=True)
if pulp.LpStatus[prob.status] == 'Optimal':
    print(output_df)

现在这将返回一个 DataFrame output_df，每个时间段和每个人都包含他们是否有人:

#                      Staffed
#Timeslot Staffmember         
#1        C1               1.0
#         C2               1.0
#         C3               1.0
#         C4               0.0
#         C5               0.0
#         C6               0.0
#         C7               0.0
#         C8               0.0
#         C9               0.0
#         C10              0.0
#         C11              0.0
#2        C1               1.0
#         C2               1.0

我修改了http://benalexkeen.com/linear-programming-with-python-and-pulp-part-5/的代码这是一个很好的 PuLP 和线性规划教程，所以一定要检查一下。

附录:您的要求不可行。

根据您的条件，这实际上将返回 'Infeasible'。很容易看出这是为什么:
您可以看到在最后几个时间段中需要的工作人员多于可用的工作人员。此图由以下人员创建:

fig, ax = plt.subplots()
staffing_requirements.plot(y='People', ax=ax, label='Required', drawstyle='steps-mid')
availability_per_member.groupby(level='Timeslot')['Available'].sum().plot(ax=ax, 
                               label='Available', drawstyle='steps-mid')
plt.legend()

关于python - 创建可用值的分布 - Python，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/55330016/

文章推荐： python - 在 Python 中获取 GIF 图像的第一帧？

文章推荐： java - Hibernate JCache 5.4.3.Final 不适用于 JCache 5.4.2.Final 配置

文章推荐： php - Cron 作业不是每 30 分钟开始一次

文章推荐： java - 如何以日志格式获取应用程序运行的端口号？

python - Python 中的集群或合并集群以减少组数 (Python)
我正在处理一组标记为 160 个组的 173k 点。我想通过合并最接近的(到 9 或 10 个组)来减少组/集群的数量。我搜索过 sklearn 或类似的库，但没有成功。我猜它只是通过 knn 聚类
python - python 列表的子集基于同一列表的元素组，pythonically
我有一个扁平数字列表，这些数字逻辑上以 3 为一组，其中每个三元组是 (number, __ignored, flag[0 or 1])，例如: [7,56,1, 8,0,0, 2,0,0, 6,1,
python - 激活 Python 虚拟环境并在另一个 Python 脚本中调用 Python 脚本
我正在使用 pipenv 来管理我的包。我想编写一个 python 脚本来调用另一个使用不同虚拟环境(VE)的 python 脚本。如何运行使用 VE1 的 python 脚本 1 并调用另一个 p
python - 在焕然一新的 Python 环境中以编程方式从 Python 内部执行 Python 文件
假设我有一个文件 script.py 位于 path = "foo/bar/script.py"。我正在寻找一种在 Python 中通过函数 execute_script() 从我的主要 Python
python - 从 python 脚本但在 python 脚本之外运行 python 脚本
这听起来像是谜语或笑话，但实际上我还没有找到这个问题的答案。问题到底是什么？我想运行 2 个脚本。在第一个脚本中，我调用另一个脚本，但我希望它们继续并行，而不是在两个单独的线程中。主要是我不希望第
python - 使用不同的 python 从 python 运行 python 脚本
我有一个带有 python 2.5.5 的软件。我想发送一个命令，该命令将在 python 2.7.5 中启动一个脚本，然后继续执行该脚本。我试过用 #!python2.7.5 和http://re
python - 为什么从 Python 命令行调用 Python 时 Python 无法找到并运行我的脚本？
我在 python 命令行(使用 python 2.7)中，并尝试运行 Python 脚本。我的操作系统是 Windows 7。我已将我的目录设置为包含我所有脚本的文件夹，使用: os.chdir("
python - 使用动态版本的 Python 执行嵌入的 Python 代码时出现致命的 Python 错误
剧透:部分解决(见最后)。以下是使用 Python 嵌入的代码示例: #include int main(int argc, char** argv) { Py_SetPythonHome
python - python 中识别 python 数组或列表中最大累积差异的最快方法是什么？
假设我有以下列表，对应于及时的股票价格: prices = [1, 3, 7, 10, 9, 8, 5, 3, 6, 8, 12, 9, 6, 10, 13, 8, 4, 11] 我想确定以下总体上最
python - (Python) 通过单选按钮 python 更新背景
所以我试图在选择某个单选按钮时更改此框架的背景。我的框架位于一个类中，并且单选按钮的功能位于该类之外。 (这样我就可以在所有其他框架上调用它们。) 问题是每当我选择单选按钮时都会出现以下错误: co
python - python 中的字符串与正则表达式比较在 python 中失败
我正在尝试将字符串与 python 中的正则表达式进行比较，如下所示， #!/usr/bin/env python3 import re str1 = "Expecting property name
python - python 如何加载Boost.Python 库？
考虑以下原型(prototype) Boost.Python 模块，该模块从单独的 C++ 头文件中引入类“D”。 /* file: a/b.cpp */ BOOST_PYTHON_MODULE(c)
python - python 检查模块 python 的问题
如何编写一个程序来“识别函数调用的行号？” python 检查模块提供了定位行号的选项，但是， def di(): return inspect.currentframe().f_back.f_l
python - 系统 python 与用户 python
我已经使用 macports 安装了 Python 2.7，并且由于我的 $PATH 变量，这就是我输入 $ python 时得到的变量。然而，virtualenv 默认使用 Python 2.6，除
python - [Python] : Python re. 长字符串行的搜索速度优化
我只想问如何加快 python 上的 re.search 速度。我有一个很长的字符串行，长度为 176861(即带有一些符号的字母数字字符)，我使用此函数测试了该行以进行研究: def getExe
python - 编辑字符串 python 正则表达式 python
list1= [u'%app%%General%%Council%', u'%people%', u'%people%%Regional%%Council%%Mandate%', u'%ppp%%Ge
python - Python 映射中的副作用(Python "do" block )
这个问题在这里已经有了答案: Is it Pythonic to use list comprehensions for just side effects? (7 个答案) 关闭 4 个月前。告
python - 使用其值逻辑组合两个 python 列表 - Python
我想用 Python 将两个列表组合成一个列表，方法如下: a = [1,1,1,2,2,2,3,3,3,3] b= ["Sun", "is", "bright", "June","and" ,"Ju
python - Boost.Python python 链接错误
我正在运行带有最新 Boost 发行版 (1.55.0) 的 Mac OS X 10.8.4 (Darwin 12.4.0)。我正在按照说明 here构建包含在我的发行版中的教程 Boost-Pyth
python - 在 Python 中仅使用内置库制作一个基本的网络抓取工具 - Python
学习 Python，我正在尝试制作一个没有任何第 3 方库的网络抓取工具，这样过程对我来说并没有简化，而且我知道我在做什么。我浏览了一些在线资源，但所有这些都让我对某些事情感到困惑。 html 看起来

太空宇宙

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - 创建可用值的分布 - Python

数据操纵为成功做好准备

用`pulp`解决整数规划问题

附录:您的要求不可行。

首页

博学

6Ren·AI

商城

python - 创建可用值的分布 - Python

数据操纵为成功做好准备

用pulp解决整数规划问题

附录:您的要求不可行。

用`pulp`解决整数规划问题