Apache ORC

ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. It is optimized for large streaming reads, but with integrated support for finding required rows quickly. Storing data in a columnar format lets the reader read, decompress, and process only the values that are required for the current query. Because ORC files are type-aware, the writer chooses the most appropriate encoding for the type and builds an internal index as the file is written. Predicate pushdown uses those indexes to determine which stripes in a file need to be read for a particular query and the row indexes can narrow the search to a particular set of 10,000 rows. ORC supports the complete set of types in Hive, including the complex types: structs, lists, maps, and unions.

ORC Format

This project includes ORC specifications and the protobuf definition. Apache ORC Format 1.0.0 is designed to be used for Apache ORC 2.0+.

Releases:

Maven Central:
Downloads: Apache ORC downloads
Release tags: Apache ORC Format releases
Plan: Apache ORC Format future release plan

The current build status:

Main branch

Bug tracking: Apache ORC Format Issues

Building

./mvnw install

Name		Name	Last commit message	Last commit date
Latest commit History 19 Commits
.github		.github
specification		specification
src/main/proto		src/main/proto
.asf.yaml		.asf.yaml
.gitignore		.gitignore
LICENSE		LICENSE
NOTICE		NOTICE
README.md		README.md
mvnw		mvnw
pom.xml		pom.xml

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Apache ORC

ORC Format

Building

About

Releases

Packages

License

cxzl25/orc-format

Folders and files

Latest commit

History

Repository files navigation

Apache ORC

ORC Format

Building

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Packages