python - Google App Engine 中不止一种数据存储类型的 MapReduce-6ren

python - Google App Engine 中不止一种数据存储类型的 MapReduce

转载作者：太空狗更新时间：2023-10-30 00:04:15

26

4

我刚看了Batch data processing with App Engine session of Google I/O 2010 , 阅读 MapReduce article from Google Research 的部分内容现在我正在考虑使用 MapReduce on Google App Engine用 Python 实现推荐系统。

我更喜欢使用 appengine-mapreduce 而不是 Task Queue API，因为前者可以轻松迭代某种类型的所有实例、自动批处理、自动任务链等。问题是:我的推荐系统需要计算实例之间的相关性两种不同的模型，即两种不同类型的实例。

例子:我有这两个模型:用户和项目。每个都有一个标签列表作为属性。以下是计算用户和项目之间相关性的函数。请注意，应该为用户和项目的每个组合调用 calculateCorrelation:

def calculateCorrelation(user, item):
    return calculateCorrelationAverage(u.tags, i.tags)

def calculateCorrelationAverage(tags1, tags2):
    correlationSum = 0.0
    for (tag1, tag2) in allCombinations(tags1, tags2):
        correlationSum += correlation(tag1, tag2)
    return correlationSum / (len(tags1) + len(tags2))

def allCombinations(list1, list2):
    combinations = []
    for x in list1:
        for y in list2:
            combinations.append((x, y))
    return combinations

但是 calculateCorrelation 在 appengine-mapreduce 中不是一个有效的 Mapper，也许这个函数甚至与 MapReduce 计算概念不兼容。然而，我需要确定...如果我拥有 appengine-mapreduce 的优势(例如自动批处理和任务链)，那真的很棒。

有什么解决办法吗？

我应该定义自己的 InputReader 吗？读取两种不同类型的所有实例的新 InputReader 是否与当前的 appengine-mapreduce 实现兼容？

或者我应该尝试以下方法吗？

将这两种类型的所有实体的所有键，两个两个地组合成一个新模型的实例(可能使用 MapReduce)
使用映射器迭代这个新模型的实例
对于每个实例，使用其中的键来获取两个不同种类的实体并计算它们之间的相关性。

最佳答案

根据 Nick Johnson 的建议，我编写了自己的 InputReader。这个阅读器从两种不同的类型中获取实体。它产生包含这些实体的所有组合的元组。在这里:

class TwoKindsInputReader(InputReader):
    _APP_PARAM = "_app"
    _KIND1_PARAM = "kind1"
    _KIND2_PARAM = "kind2"
    MAPPER_PARAMS = "mapper_params"

    def __init__(self, reader1, reader2):
        self._reader1 = reader1
        self._reader2 = reader2

    def __iter__(self):
        for u in self._reader1:
            for e in self._reader2:
                yield (u, e)

    @classmethod
    def from_json(cls, input_shard_state):
        reader1 = DatastoreInputReader.from_json(input_shard_state[cls._KIND1_PARAM])
        reader2 = DatastoreInputReader.from_json(input_shard_state[cls._KIND2_PARAM])

        return cls(reader1, reader2)

    def to_json(self):
        json_dict = {}
        json_dict[self._KIND1_PARAM] = self._reader1.to_json()
        json_dict[self._KIND2_PARAM] = self._reader2.to_json()
        return json_dict

    @classmethod
    def split_input(cls, mapper_spec):
        params = mapper_spec.params
        app = params.get(cls._APP_PARAM)
        kind1 = params.get(cls._KIND1_PARAM)
        kind2 = params.get(cls._KIND2_PARAM)
        shard_count = mapper_spec.shard_count
        shard_count_sqrt = int(math.sqrt(shard_count))

        splitted1 = DatastoreInputReader._split_input_from_params(app, kind1, params, shard_count_sqrt)
        splitted2 = DatastoreInputReader._split_input_from_params(app, kind2, params, shard_count_sqrt)
        inputs = []

        for u in splitted1:
            for e in splitted2:
                inputs.append(TwoKindsInputReader(u, e))

        #mapper_spec.shard_count = len(inputs) #uncomment this in case of "Incorrect number of shard states" (at line 408 in handlers.py)
        return inputs

    @classmethod
    def validate(cls, mapper_spec):
        return True #TODO

当您需要处理两种实体的所有组合时，应使用此代码。您也可以将其推广到两种以上。

这是 TwoKindsInputReader 的有效 mapreduce.yaml:

mapreduce:
- name: recommendationMapReduce
  mapper:
    input_reader: customInputReaders.TwoKindsInputReader
    handler: recommendation.calculateCorrelationHandler
    params:
    - name: kind1
      default: kinds.User
    - name: kind2
      default: kinds.Item
    - name: shard_count
      default: 16

关于python - Google App Engine 中不止一种数据存储类型的 MapReduce，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/3766154/

26

4

0

文章推荐： python - 我怎么知道子进程何时死亡？

文章推荐： c# - 什么样的内存语义支配 C# 中的数组赋值？

文章推荐： c# - 如何修补 .NET 程序集？

google-app-engine - Google Cloud 中的 Google Compute Engine、App Engine 和 Container Engine 有什么区别？
Google Cloud Compute 中的 Google Compute Engine、App Engine 和 Container Engine 之间的实际区别是什么？什么时候使用什么？有什么
google-app-engine - 用于 App Engine 网址提取的 Compute Engine 防火墙
我有一个在 Google App Engine 中运行的应用程序，它访问在 Google Compute Engine 中的机器上运行的服务。 Google App Engine 应用程序是该服务唯一
google-app-engine - Google App Engine 通过内部网络与 Compute Engine 通信
我们正在谷歌云中构建一个应用程序。我们使用 App Engine 作为前端，使用 Compute Engine 作为后端。在这些 Compute Engine 实例上，我正在运行一个接受特定“命令”消
google-app-engine - 允许 App Engine 应用程序访问另一个 App Engine 应用程序的数据存储区
我有一个现有的 GAE 应用程序(我们称之为应用程序 A)正在运行的情况，但由于非技术原因无法修改。当用户迁移到新的客户端版本时，我们需要将他们的数据从应用程序 A 迁移到新的 GAE 应用程序(我称
google-app-engine - Google App Engine 模块主机名 : not an App Engine context
我正在尝试发现 App Engine 上的其他已部署服务。类似于 this文章建议。我的代码是这样的: import ( "fmt" "net/http" "google.g
google-app-engine - 'Google App Engine' 比 'Google Compute Engine' 贵得多吗？
我想在我的网站上为“图像处理”事件设置服务器。如果我在 GCE 中使用“n1-standard-1”实例，GAE 中的可比功率是多少？是因为我算错了，还是同一个功率两者价格相差很大？最佳答案按小时
google-app-engine - 连接 Google App Engine 和 Google Compute Engine
我在 Googl Compute Engine 和 Google App Engine 标准环境中的应用程序中创建了一个 VM 实例。我打算在 App Engine 中使用我的应用程序，在 Compu
google-app-engine - 无法使用 App Engine SDK 在 App Engine 上部署
我像往常一样使用 appcfg.py 更新我的应用程序，但收到一条错误消息。我试过 appcfg.py 回滚，两次尝试之间等了十分钟，但我仍然收到相同的错误消息。我该怎么办？无法对 apps/dev
google-app-engine - 如何使用/Google Compute Engine 安全地配置 App Engine 套接字
我想在 Google Compute Engine 上放置一个 Redis 服务器，并通过 AppEngine 的套接字支持与其对话。唯一的问题是似乎没有特定的防火墙规则说“此 AppEngine 应
google-app-engine - Google App Engine 和 Google Compute Engine 有什么区别？
我想知道 App Engine 和 Compute Engine 之间有什么区别。任何人都可以向我解释其中的区别吗？最佳答案 App Engine 是一种平台即服务。这意味着您只需部署代码，平台会为
google-app-engine - 如何管理 App Engine Go 运行时上下文以避免 App Engine 锁定？
我正在编写一个在 App Engine 的 Go 运行时上运行的 Go 应用程序。我注意到几乎所有使用 App Engine 服务(例如 Datastore、Mail 甚至 Capabilities
process - 在 Grid Engine/Sun Grid Engine/Son of Grid Engine 上使用 Docker
是否有人有在 Grid Engine/Sun Grid Engine/Son of Grid Engine 上运行 Docker 的经验，并且能够 monitor the resource used
google-app-engine - 我可以使用 intellj app-engine 插件将我的 grails 应用程序部署到 app-engine
我读了很多论坛，因为 grails app-engine 插件多年来没有更新，所以不可能将 grails 应用程序部署到谷歌应用程序引擎。当我准备放弃时，我发现使用 intellij 部署项目是可能的
google-app-engine - 仅从 Google App Engine 访问 Google Compute Engine 实例？
当前设置，运行 Windows Server 2012 (GCE Server 2012) 的谷歌计算引擎运行 Debian Wheezy(GCE 服务器 Wheezy)的 Google 计算引擎
google-app-engine - Google App Engine Flexible 和 Google Container Engine 之间的区别？
特定于基于 Docker 的部署，这两者之间有什么区别？由于 Google App Engine Flexible 现在也支持基于 Dockerfile 的部署，并且它也是完全托管的服务，因此它似乎比
google-app-engine - Google Kubernetes Engine (GKE) 和 Google Compute Engine (GCE) 在服务器管理方面有什么区别？
我相信 Google Kubernetes Engine (GKE) 在 Google Compute Engine (GCE) 上运行。那么，在服务器管理方面使用 Google Kubernetes
google-app-engine - App Engine 和 Compute Engine 实例之间是否存在 AWS "security groups"的等效项？
TLDR；关于这个问题有任何更新吗？ Google App Engine communicate with Compute Engine over internal network -- 是否可以在同
google-app-engine - 正确的 GOPATH 包含来自 App Engine SDK 的 App Engine 库？
我正在尝试使用 Go SDK 为 App Engine 编写应用程序，但它似乎与单元测试有一种有趣的关系。人有written libraries左右this original, outdated一组工
google-app-engine - Google App Engine 可以在没有外部 IP 的情况下向同一项目中的 Compute Engine 实例发出 http 请求吗？
在 App Engine 中，我想对在同一个 Google 云项目中创建的 Compute Engine 实例上运行的网络服务器进行 http fetch 调用，我想知道是否可以在不启用的情况下对实例
google-app-engine - Cloud Datastore 客户端库与 App Engine Go Standard 上的 App Engine SDK
在编写 Go App Engine 标准应用程序时，过去的情况是您必须使用 App Engine SDK访问数据存储。然而，最近(从 Go 1.11 开始？)，如果你只使用 Cloud Datasto

首页

博学

6Ren·AI

商城

python - Google App Engine 中不止一种数据存储类型的 MapReduce