A cross-platform (Windows, MAC, Linux) desktop application to view common bigdata binary format like Parquet, ORC, Avro, etc. Support local file system, HDFS, AWS S3, etc. Add basic data analysis functions like aggregate operations and checking data proportions.
Note, you're recommended to download release v1.1.1 to if you just want to view local bigdata binary files, it's lightweight without dependency to AWS SDK, Azure SDK, etc. Quite honestly, you can download data files from web portal of AWS, Azure ,etc. before viewing it with this tool. The reason why I integrated the cloud storage system's SDK into this tool is more like a demo of how to use Java to read files from specific storage system.
Build
section to build from source code.java -jar BigdataFileViewer-1.3-SNAPSHOT-jar-with-dependencies.jar
or invoke with parameter '-a' which means enable experimental analytics featureInput maximum row number
-> "Go"Check schema information by unfolding "Schema Information" panel
mvn package
to build an all-in-one runnable jar. If you're using JDK11 or higher, using mvn package -Pjava11
to include openJFX as Maven dependency directly. If you're using Mac with M1 chip, please make sure your JDK version is 17+ and please build with mvn package -Pjava17
If you're using Java 1.8 or lower, make sure the Java has javafx bound. For example, I installed openjdk 1.8 on Ubuntu 18.04 and it has no javafx bound, then I installed it following guide here.
The INT96 data type is deprecated per parqeut-mr, so please expect java.lang.IllegalArgumentException: INT96 is deprecated.
if you're trying to open a parquet file contains INT96 data type.
Speicial thanks to Meindert Deen, sedzisz, barabulkit, marcomalva who have contributed to the project.