gpt4 book ai didi

python - 多线程 Python 应用程序和套接字连接的问题

转载 作者:太空狗 更新时间:2023-10-29 20:44:20 25 4
gpt4 key购买 nike

我正在调查在具有 4G RAM 的 Ubuntu 计算机上运行的 Python 应用程序的问题。该工具将用于审核服务器(我们更愿意推出自己的工具)。它使用线程连接到大量服务器,许多 TCP 连接失败。但是,如果我在启动每个线程之间添加 1 秒的延迟,那么大多数连接都会成功。我使用这个简单的脚本来调查可能发生的情况:

#!/usr/bin/python

import sys
import socket
import threading
import time

class Scanner(threading.Thread):
def __init__(self, host, port):
threading.Thread.__init__(self)
self.host = host
self.port = port
self.status = ""

def run(self):
self.sk = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
self.sk.settimeout(20)
try:
self.sk.connect((self.host, self.port))
except Exception, err:
self.status = str(err)
else:
self.status = "connected"
finally:
self.sk.close()


def get_hostnames_list(filename):
return open(filename).read().splitlines()

if (__name__ == "__main__"):
hostnames_file = sys.argv[1]
hosts_list = get_hostnames_list(hostnames_file)
threads = []
for host in hosts_list:
#time.sleep(1)
thread = Scanner(host, 443)
threads.append(thread)
thread.start()

for thread in threads:
thread.join()
print "Host: ", thread.host, " : ", thread.status

如果我在 time.sleep(1) 运行此命令时针对 300 台主机进行了注释,则许多连接会因超时错误而失败,而如果我延迟一秒,它们不会超时。我是否在功能更强大的机器上运行的另一个 Linux 发行版上尝试过该应用程序,并且没有那么多连接错误?是因为内核限制吗?我可以做些什么来使连接正常工作而不会造成延迟?

更新

我还尝试了一个限制池中可用线程数的程序。通过将其减少到 20,我可以让所有连接正常工作,但它每秒只检查大约 1 个主机。因此,无论我尝试什么(进入 sleep(1) 或限制并发线程数),我似乎每秒都无法检查超过 1 个主机。

更新

我刚找到这个 question 这似乎与我所看到的相似。

更新

我想知道使用 twisted 编写此代码是否有帮助?任何人都可以展示我的示例使用扭曲编写的样子吗?

最佳答案

你可以试试 gevent :

from gevent.pool import Pool    
from gevent import monkey; monkey.patch_all() # patches stdlib
import sys
import logging
from httplib import HTTPSConnection
from timeit import default_timer as timer
info = logging.getLogger().info

def connect(hostname):
info("connecting %s", hostname)
h = HTTPSConnection(hostname, timeout=2)
try: h.connect()
except IOError, e:
info("error %s reason: %s", hostname, e)
else:
info("done %s", hostname)
finally:
h.close()

def main():
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(message)s")
info("getting hostname list")
hosts_file = sys.argv[1] if len(sys.argv) > 1 else "hosts.txt"
hosts_list = open(hosts_file).read().splitlines()
info("spawning jobs")
pool = Pool(20) # limit number of concurrent connections
start = timer()
for _ in pool.imap(connect, hosts_list):
pass
info("%d hosts took us %.2g seconds", len(hosts_list), timer() - start)

if __name__=="__main__":
main()

它每秒可以处理多个主机。

输出

2011-01-31 11:08:29,052 getting hostname list
2011-01-31 11:08:29,052 spawning jobs
2011-01-31 11:08:29,053 connecting www.yahoo.com
2011-01-31 11:08:29,053 connecting www.abc.com
2011-01-31 11:08:29,053 connecting www.google.com
2011-01-31 11:08:29,053 connecting stackoverflow.com
2011-01-31 11:08:29,053 connecting facebook.com
2011-01-31 11:08:29,054 connecting youtube.com
2011-01-31 11:08:29,054 connecting live.com
2011-01-31 11:08:29,054 connecting baidu.com
2011-01-31 11:08:29,054 connecting wikipedia.org
2011-01-31 11:08:29,054 connecting blogspot.com
2011-01-31 11:08:29,054 connecting qq.com
2011-01-31 11:08:29,055 connecting twitter.com
2011-01-31 11:08:29,055 connecting msn.com
2011-01-31 11:08:29,055 connecting yahoo.co.jp
2011-01-31 11:08:29,055 connecting taobao.com
2011-01-31 11:08:29,055 connecting google.co.in
2011-01-31 11:08:29,056 connecting sina.com.cn
2011-01-31 11:08:29,056 connecting amazon.com
2011-01-31 11:08:29,056 connecting google.de
2011-01-31 11:08:29,056 connecting google.com.hk
2011-01-31 11:08:29,188 done www.google.com
2011-01-31 11:08:29,189 done google.com.hk
2011-01-31 11:08:29,224 error wikipedia.org reason: [Errno 111] Connection refused
2011-01-31 11:08:29,225 done google.co.in
2011-01-31 11:08:29,227 error msn.com reason: [Errno 111] Connection refused
2011-01-31 11:08:29,228 error live.com reason: [Errno 111] Connection refused
2011-01-31 11:08:29,250 done google.de
2011-01-31 11:08:29,262 done blogspot.com
2011-01-31 11:08:29,271 error www.abc.com reason: [Errno 111] Connection refused
2011-01-31 11:08:29,465 done amazon.com
2011-01-31 11:08:29,467 error sina.com.cn reason: [Errno 111] Connection refused
2011-01-31 11:08:29,496 done www.yahoo.com
2011-01-31 11:08:29,521 done stackoverflow.com
2011-01-31 11:08:29,606 done youtube.com
2011-01-31 11:08:29,939 done twitter.com
2011-01-31 11:08:33,056 error qq.com reason: timed out
2011-01-31 11:08:33,057 error taobao.com reason: timed out
2011-01-31 11:08:33,057 error yahoo.co.jp reason: timed out
2011-01-31 11:08:34,466 done facebook.com
2011-01-31 11:08:35,056 error baidu.com reason: timed out
2011-01-31 11:08:35,057 20 hosts took us 6 seconds

关于python - 多线程 Python 应用程序和套接字连接的问题,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/4783735/

25 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com