Enum 的 Avro Schema Evolution – 反序列化崩溃-6ren

Enum 的 Avro Schema Evolution – 反序列化崩溃

转载作者：行者123 更新时间：2023-12-03 20:51:27

25

4

我在两个单独的 AVCS 模式文件中定义了记录的两个版本。我用命名空间来区分版本
SimpleV1.avsc

{
  "type" : "record",
  "name" : "Simple",
  "namespace" : "test.simple.v1",
  "fields" : [ 
      {
        "name" : "name",
        "type" : "string"
      }, 
      {
        "name" : "status",
        "type" : {
          "type" : "enum",
          "name" : "Status",
          "symbols" : [ "ON", "OFF" ]
        },
        "default" : "ON"
      }
   ]
}

示例 JSON

{"name":"A","status":"ON"}

版本 2 只有一个带有默认值的附加说明字段。
SimpleV2.avsc

{
  "type" : "record",
  "name" : "Simple",
  "namespace" : "test.simple.v2",
  "fields" : [ 
      {
        "name" : "name",
        "type" : "string"
      }, 
      {
        "name" : "description",
        "type" : "string",
        "default" : ""
      }, 
      {
        "name" : "status",
        "type" : {
          "type" : "enum",
          "name" : "Status",
          "symbols" : [ "ON", "OFF" ]
        },
        "default" : "ON"
      }
   ]
}

示例 JSON

{"name":"B","description":"b","status":"ON"}

两种模式都被序列化为 Java 类。
在我的示例中，我将测试向后兼容性。由 V1 写入的记录应由使用 V2 的阅读器读取。我想看到插入了默认值。只要我不使用枚举，这就是有效的。

public class EnumEvolutionExample {

    public static void main(String[] args) throws IOException {
        Schema schemaV1 = new org.apache.avro.Schema.Parser().parse(new File("./src/main/resources/SimpleV1.avsc"));
        //works as well
        //Schema schemaV1 = test.simple.v1.Simple.getClassSchema();
        Schema schemaV2 = new org.apache.avro.Schema.Parser().parse(new File("./src/main/resources/SimpleV2.avsc"));

        test.simple.v1.Simple simpleV1 = test.simple.v1.Simple.newBuilder()
                .setName("A")
                .setStatus(test.simple.v1.Status.ON)
                .build();
        
        
        SchemaPairCompatibility schemaCompatibility = SchemaCompatibility.checkReaderWriterCompatibility(
                schemaV2,
                schemaV1);
        //Checks that writing v1 and reading v2 schemas is compatible
        Assert.assertEquals(SchemaCompatibilityType.COMPATIBLE, schemaCompatibility.getType());
        
        byte[] binaryV1 = serealizeBinary(simpleV1);
        
        //Crashes with: AvroTypeException: Found test.simple.v1.Status, expecting test.simple.v2.Status
        test.simple.v2.Simple v2 = deSerealizeBinary(binaryV1, new test.simple.v2.Simple(), schemaV1);
        
    }
    
    public static byte[] serealizeBinary(SpecificRecord record) {
        DatumWriter<SpecificRecord> writer = new SpecificDatumWriter<>(record.getSchema());
        byte[] data = new byte[0];
        ByteArrayOutputStream stream = new ByteArrayOutputStream();
        Encoder binaryEncoder = EncoderFactory.get()
            .binaryEncoder(stream, null);
        try {
            writer.write(record, binaryEncoder);
            binaryEncoder.flush();
            data = stream.toByteArray();
        } catch (IOException e) {
            System.out.println("Serialization error " + e.getMessage());
        }

        return data;
    }
    
    public static <T extends SpecificRecord> T deSerealizeBinary(byte[] data, T reuse, Schema writer) {
        Decoder decoder = DecoderFactory.get().binaryDecoder(data, null);
        DatumReader<T> datumReader = new SpecificDatumReader<>(writer, reuse.getSchema());
        try {
            T datum = datumReader.read(null, decoder);
            return datum;
        } catch (IOException e) {
            System.out.println("Deserialization error" + e.getMessage());
        }
        return null;
    }

}

checkReaderWriterCompatibility 方法确认模式是兼容的。
但是当我反序列化时，我收到以下异常

Exception in thread "main" org.apache.avro.AvroTypeException: Found test.simple.v1.Status, expecting test.simple.v2.Status
    at org.apache.avro.io.ResolvingDecoder.doAction(ResolvingDecoder.java:309)
    at org.apache.avro.io.parsing.Parser.advance(Parser.java:86)
    at org.apache.avro.io.ResolvingDecoder.readEnum(ResolvingDecoder.java:260)
    at org.apache.avro.generic.GenericDatumReader.readEnum(GenericDatumReader.java:267)
    at org.apache.avro.generic.GenericDatumReader.readWithoutConversion(GenericDatumReader.java:181)
    at org.apache.avro.specific.SpecificDatumReader.readField(SpecificDatumReader.java:136)
    at org.apache.avro.generic.GenericDatumReader.readRecord(GenericDatumReader.java:247)
    at org.apache.avro.specific.SpecificDatumReader.readRecord(SpecificDatumReader.java:123)
    at org.apache.avro.generic.GenericDatumReader.readWithoutConversion(GenericDatumReader.java:179)
    at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:160)
    at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:153)
    at test.EnumEvolutionExample.deSerealizeBinary(EnumEvolutionExample.java:70)
    at test.EnumEvolutionExample.main(EnumEvolutionExample.java:45)

我不明白为什么 Avro 认为它有一个 v1.Status。命名空间不是编码的一部分。
这是一个错误还是有人知道如何运行它？

最佳答案

找到了解决方法。我将枚举移动到“未版本化”命名空间。所以它在两个版本中都是一样的。
但实际上它对我来说似乎是一个错误。转换记录不是问题，但枚举不起作用。两者都是 Avro 中的复杂类型。

{
  "type" : "record",
  "name" : "Simple",
  "namespace" : "test.simple.v1",
  "fields" : [ 
      {
        "name" : "name",
        "type" : "string"
      }, 
      {
        "name" : "status",
        "type" : {
          "type" : "enum",
          "name" : "Status",
          "namespace" : "test.model.unversioned",
          "symbols" : [ "ON", "OFF" ]
        },
        "default" : "ON"
      }
   ]
}

关于Enum 的 Avro Schema Evolution – 反序列化崩溃，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/62596990/

25

4

0

文章推荐： animation - 全屏封面/模态的替代动画 - iOS 14

avro - 如何在不再次定义的情况下在另一种 Avro 类型中使用 Avro 类型？
我在名为 commonSourceMetadata.avsc 的 json 文件中定义了一个名为 "some.package.SourceMetadata" 的 Avro 类型: { "type"
avro - Avro 中特定数据类型的最佳实践
我很想了解在 Avro 中编码两种非常特定类型的数据的最佳实践:时间戳和 IP 地址。我遇到了时间戳 ( https://issues.apache.org/jira/browse/AVRO-739
avro - 最大尺寸限制在 avro
如何在 Avro Schema 生成中为数据类型设置最大大小/长度限制。例如:在模式中，我想指定一个字段，该字段采用最大 len 25 的字符串。最佳答案我相信您可以使用“固定”avro 类型并指
avro - Avro 是否支持必填字段？
即是否可以使字段需要类似于 ProtoBuf: 消息搜索请求{ 需要字符串查询 = 1; } 最佳答案默认情况下，Avro 中的所有字段都是必需的。照原样 mentioned在官方文档中，如果你想
hadoop - Flume:Directory to Avro -> Avro to HDFS - Not valid avro after transfer
我有用户编写 AVRO 文件，我想使用 Flume 将所有这些文件移动到使用 Flume 的 HDFS 中。所以我以后可以使用 Hive 或 Pig 来查询/分析数据。在客户端我安装了 flume
avro - 为具有多种记录类型的数组创建 avro 模式？
我正在为似乎具有多个对象数组的 JSON 有效负载创建 avro 模式。我不确定如何在模式中表示这一点。有问题的关键是 content: { "id": "channel-id", "name
avro - 您可以将数据附加到现有的 Avro 数据文件中吗？
似乎没有任何方法可以将数据附加到现有的 Avro 序列化文件中。我想让多个进程写入一个 avro 文件，但看起来每次打开它时，我都会从头开始。我不想读入所有数据，然后再将其写回。使用 ruby
avro - Apache Avro 架构示例和文档
我试图定义一个不太平凡的 Avro 模式，但收效甚微；当它不会抛出架构语法错误时，它不会生成我试图在架构中定义的所有类型。是否有 avsc 定义的可能内容的完整规范？我一直根据我从 Doc 规范中理
Avro-Tools JSON 到 Avro 架构失败 : org. apache.avro.SchemaParseException:未定义名称:
我正在尝试使用 avro-tools-1.7.4.jar create schema 命令创建两个 Avro 模式。我有两个 JSON 模式，如下所示: { "name": "TestAvro",
hadoop - 将表的属性从 avro.schema.literal 设置为 avro.schema.url 后，Hive avro 表架构未更新
首先，我创建了一个如下所示的 avro hive 表。 CREATE EXTERNAL TABLE user STORED AS AVRO LOCATION '/work/user' TBLPROPE
hadoop - Avro 序列化和 Avro 格式的区别
我正在读一本书 Hadoop application architectures，这本书很老但很有趣，在阅读时，我注意到 Avro 被认为是数据序列化框架，而 Parquet 被认为是列数据格式。我
avro - 到目前为止，Apache Avro 中代码日期字段的最佳实践是什么？
我一直在四处寻找，看到了 jira https://issues.apache.org/jira/browse/AVRO-739对于这个问题，但我对用户文档中的日期时间的 avro 支持没有更好的了解
scala - 为什么在我使用 com.databricks.spark.avro 时必须添加 org.apache.spark.avro 依赖才能在 Spark2.4 中读/写 avro 文件？
我尝试在安装了 Spark 2.4.8 的 Cloud Dataproc 集群 1.4 上运行我的 Spark/Scala 代码 2.3.0。我在读取 avro 文件时遇到错误。这是我的代码: spa
avro - 如何在 Avro 中将记录与 map 混合？
我正在处理 JSON 格式的服务器日志，我想以 Parquet 格式将我的日志存储在 AWS S3 上(并且 Parquet 需要 Avro 模式)。首先，所有日志都有一组共同的字段，其次，所有日志都
java - 为什么 avro 无法从 .avro 文件中获取架构？
这是来自教程点的解串器。 public class Deserialize { public static void main(String args[]) throws Exception{
avro - 为什么 avro 生成的 java 代码有这么多不推荐使用的字段
我正在使用 avro-maven-plugin 1.8.1 从 schema 生成 java 代码，所有字段都是公共(public)的且已弃用，如下所示: public class data_el
avro - 在 Avro IDL 中，如何导入外部提供的架构？
一个简单的例子说明了我的问题。本质上，我正在处理一个跨多个存储库拆分代码的大型项目。在 repo 1 中，在 .avdl 文件中定义了一个 Avro 模式“S1”，该文件被编译到其 Avro 生成的
c - 通过套接字发送 avro(avro c) 编码数据
通过套接字发送avro(avro c)编码数据我正在尝试将 avro 编码数据转换为字节数组(使用 memcpy)后通过套接字发送。我所做的如下所示 /客户端:client.c/ avro_datum
java - 如何在压缩的 avro 文件中获取每个 avro 记录的开始和结束？
我的问题是这样的。我有一个 2GB 的压缩 avro 文件，HDFS 上存储了大约 1000 条 avro 记录。我知道我可以编写代码来“打开这个 avro 文件”并打印出每条 avro 记录。我的问
java - Kafka Avro 序列化器和反序列化器异常。 Avro 支持的类型
我看到以下错误 exception Unsupported Avro type. Supported types are null, Boolean, Integer, Long, Float, Do

首页

博学

6Ren·AI

商城

Enum 的 Avro Schema Evolution – 反序列化崩溃