python - 使用 aiohttp/asyncio 发出 100 万次请求

python - 使用 aiohttp/asyncio 发出 100 万次请求 - 字面意思

转载作者：太空狗更新时间：2023-10-30 01:20:31

25

4

我跟进了本教程:https://pawelmhm.github.io/asyncio/python/aiohttp/2016/04/22/asyncio-aiohttp.html当我处理 50 000 个请求时，一切正常。但是我需要进行 100 万次 API 调用，然后我遇到了这段代码的问题:

    url = "http://some_url.com/?id={}"
    tasks = set()

    sem = asyncio.Semaphore(MAX_SIM_CONNS)
    for i in range(1, LAST_ID + 1):
        task = asyncio.ensure_future(bound_fetch(sem, url.format(i)))
        tasks.add(task)

    responses = asyncio.gather(*tasks)
    return await responses

因为 Python 需要创建 100 万个任务，它基本上只是滞后，然后在终端打印 Killed 消息。有什么方法可以使用预制的 url 集(或列表)的生成器？谢谢。

最佳答案

一次安排所有 100 万个任务

这是您正在谈论的代码。它最多占用 3 GB RAM，因此如果您的可用内存不足，它很可能会被操作系统终止。

import asyncio
from aiohttp import ClientSession

MAX_SIM_CONNS = 50
LAST_ID = 10**6

async def fetch(url, session):
    async with session.get(url) as response:
        return await response.read()

async def bound_fetch(sem, url, session):
    async with sem:
        await fetch(url, session)

async def fetch_all():
    url = "http://localhost:8080/?id={}"
    tasks = set()
    async with ClientSession() as session:
        sem = asyncio.Semaphore(MAX_SIM_CONNS)
        for i in range(1, LAST_ID + 1):
            task = asyncio.create_task(bound_fetch(sem, url.format(i), session))
            tasks.add(task)
        return await asyncio.gather(*tasks)

if __name__ == '__main__':
    asyncio.run(fetch_all())

使用队列简化工作

这是我的建议如何使用 asyncio.Queue将 URL 传递给工作任务。队列按需填充，没有预制的 URL 列表。

它只需要 30 MB RAM :)

import asyncio
from aiohttp import ClientSession

MAX_SIM_CONNS = 50
LAST_ID = 10**6

async def fetch(url, session):
    async with session.get(url) as response:
        return await response.read()

async def fetch_worker(url_queue):
    async with ClientSession() as session:
        while True:
            url = await url_queue.get()
            try:
                if url is None:
                    # all work is done
                    return
                response = await fetch(url, session)
                # ...do something with the response
            finally:
                url_queue.task_done()
                # calling task_done() is necessary for the url_queue.join() to work correctly

async def fetch_all():
    url = "http://localhost:8080/?id={}"
    url_queue = asyncio.Queue(maxsize=100)
    worker_tasks = []
    for i in range(MAX_SIM_CONNS):
        wt = asyncio.create_task(fetch_worker(url_queue))
        worker_tasks.append(wt)
    for i in range(1, LAST_ID + 1):
        await url_queue.put(url.format(i))
    for i in range(MAX_SIM_CONNS):
        # tell the workers that the work is done
        await url_queue.put(None)
    await url_queue.join()
    await asyncio.gather(*worker_tasks)

if __name__ == '__main__':
    asyncio.run(fetch_all())

关于python - 使用 aiohttp/asyncio 发出 100 万次请求 - 字面意思，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/38831322/

25

4

0

文章推荐： python - 多对多关系不存在。它在不同的模式中

文章推荐： c# - Inno Setup 循环遍历文件并注册每个 .NET dll

文章推荐： python - Django:如何对更新 View /表单进行单元测试

python-asyncio - Python asyncio - 增加Semaphore的值
我正在我的一个项目中使用 aiohttp 并想限制每秒发出的请求数。我正在使用 asyncio.Semaphore 来做到这一点。我的挑战是我可能想要增加/减少每秒允许的请求数。例如: limit
python-asyncio - 在 asyncio 中混合异步上下文管理器和直接等待
如何混合 async with api.open() as o: ... 和 o = await api.open() 在一个功能中？自从第一次需要带有 __aenter__ 的对象以来和
python-asyncio - 使用 asyncio 做多项终极工作
有 2 个工作:“wash_clothes”(job1) 和“setup_cleaning_robot”(job2)，每个工作需要你 7 和 3 秒，你必须做到世界末日。这是我的代码: import
python-asyncio - 如何为 asyncio 任务设置名称？
我们有一种设置线程名称的方法:thread = threading.Thread(name='Very important thread', target=foo)，然后在格式化程序中使用 %(thr
python - 使用 asyncio 生成器和 asyncio.as_completed
我有一些代码，用于抓取 URL、解析信息，然后使用 SQLAlchemy 将其放入数据库中。我尝试异步执行此操作，同时限制同时请求的最大数量。这是我的代码: async def get_url(ai
Python Asyncio 未使用 asyncio.run_coroutine_threadsafe 运行新协程
1>Python Asyncio 未使用 asyncio.run_coroutine_threadsafe 运行新的协程。下面是在Mac上进行的代码测试。 ——————————————————————
python - Asyncio.gather 与 asyncio.wait
asyncio.gather和 asyncio.wait似乎有类似的用途:我有一堆我想要执行/等待的异步事情(不一定要在下一个开始之前等待一个完成)。它们使用不同的语法，并且在某些细节上有所不同，但对
python-asyncio - 属性错误 : module 'asyncio' has no attribute 'run'
我正在尝试使用 asyncio 运行以下程序: import asyncio async def main(): print('Hello') await asyncio.sleep(
python-asyncio - 如何使用 asyncio 接口(interface)阻塞和非阻塞代码
我正在尝试在事件循环之外使用协程函数。 (在这种情况下，我想在 Django 中调用一个也可以在事件循环中使用的函数) 如果不使调用函数成为协程，似乎没有办法做到这一点。我意识到 Django 是为
python - 在 asyncio.gather 中内联链 asyncio 协程
我有一个假设 asyncio.gather设想: await asyncio.gather( cor1, [cor2, cor3], cor4, ) 我要 cor2和 cor3
Python3 和 asyncio : how to implement websocket server as asyncio instance?
我有多个服务器，每个服务器都是 asyncio.start_server 返回的实例。我需要我的 web_server 与 websockets 一起使用，以便能够使用我的 javascript 客户
Python 3 asyncio - yield from vs asyncio.async 堆栈使用
我正在使用 Python 3 asyncio 框架评估定期执行的不同模式(为简洁起见省略了实际 sleep /延迟)，我有两段代码表现不同，我无法解释原因。第一个版本使用 yield from 递归调
loop.create_task 和 asyncio.run_coroutine_threadsafe 之间的 Python asyncio 区别
从事件线程外部将协程推送到事件线程的 pythonic 方法是什么？最佳答案更新信息: 从Python 3.7 高级函数asyncio.create_task(coro)开始was added并且
python-asyncio - 如何 asyncio.gather block 中的任务+使用具有 TCP 连接限制的信号量？
我有一个大型 (1M) 数据库结果集，我想为其每一行调用一个 REST API。 API 可以接受批处理请求，但我不确定如何分割 rows 生成器，以便每个任务处理一个行列表，比如 10。我宁愿不预先
python - 混合 asyncio 和 Kivy : How to start the asyncio loop and the Kivy application at the same time?
迷失在异步中。我同时在学习Kivy和asyncio，卡在了解决运行Kivy和运行asyncio循环的问题上，无论怎么转，都是阻塞调用，需要顺序执行(好吧，我希望我是错的)，例如 loop = asy
python - asyncio python 3.6 代码到 asyncio python 3.4 代码
我有这个 3.6 异步代码: async def send(command,userPath,token): async with websockets.connect('wss://127.
python - 使用 asyncio.wait_for 和 asyncio.Semaphore 时如何正确捕获 concurrent.futures._base.TimeoutError ？
首先，我需要警告你:我是 asyncio 的新手，而且我是我马上警告你，我是 asyncio 的新手，我很难想象引擎盖下的库里有什么。这是我的代码: import asyncio semaphor
python - 当 asyncio.PriorityQueue 处于 maxsize 并且我 put() 新项目时，如何将项目从 asyncio.PriorityQueue 中推出？
我有一个asyncio.PriorityQueue，用作网络爬虫的URL队列，当我调用url_queue.get时，得分最低的URL首先从队列中删除()。当队列达到 maxsize 项时，默认行为是阻
python - 在具有 asyncio.coroutine 方法的类外部声明的 asyncio event_loop 失败并显示 "AttributeError: ' NoneType' 对象没有属性 'select'“
探索 Python 3.4.0 的 asyncio 模块，我试图创建一个类，其中包含从类外部的 event_loop 调用的 asyncio.coroutine 方法。我的工作代码如下。 impor
python-3.5 - python 3 asyncio : coroutines execution order using run_until_complete(asyncio. 等待(corutines_list))
我有一个可能是无用的问题，但尽管如此，我还是觉得我错过了一些对于理解 asyncio 的工作方式可能很重要的东西。我刚刚开始熟悉 asyncio 并编写了这段非常基本的代码: import asyn

首页

博学

6Ren·AI

商城

python - 使用 aiohttp/asyncio 发出 100 万次请求 - 字面意思

一次安排所有 100 万个任务

使用队列简化工作