If we deploy both Apache Ozone and Apache Spark on kubernetes, is it possible to achieve data locality? Or will data always have to be shuffled upon read?
How to achieve data locality with Spark and Apache Ozone? Is it possible?
95 Views Asked by Pavel Orekhov At
1
There are 1 best solutions below
Related Questions in APACHE-SPARK
- Getting error while running spark-shell on my system; pyspark is running fine
- ingesting high volume small size files in azure databricks
- Spark load all partions at once
- Databricks Delta table / Compute job
- Autocomplete not working for apache spark in java vscode
- How to overwrite a single partition in Snowflake when using Spark connector
- Parse multiple record type fixedlength file with beanio gives oom and timeout error for 10GB data file
- includeExistingFiles: false does not work in Databricks Autoloader
- Spark connectors from Azure Databricks to Snowflake using AzureAD login
- SparkException: Task failed while writing rows, caused by Futures timed out
- Configuring Apache Spark's MemoryStream to simulate Kafka stream
- Databricks can't find a csv file inside a wheel I installed when running from a Databricks Notebook
- Add unique id to rows in batches in Pyspark dataframe
- Does Spark Dynamic Allocation depend on external shuffle service to work well?
- Does Spark structured streaming support chained flatMapGroupsWithState by different key?
Related Questions in OZONE
- How can I connect my SpringBoot application to Apache Ozone using Kerberos
- Breakpoint in HardFault_Handler doesn't work
- How to achieve data locality with Spark and Apache Ozone? Is it possible?
- Calculate SOMO35 from netcdf file
- Apache Ozone Java API: PERMISSION_DENIED for read and create keys of file-type
- I fire a create table statement on hive catalog from trino, the query gets stuck
- How to access data from ozone using flink?
- a tool or library that can be controlled by the .NET Framework Library (.DLL) to edit an .elf file
- How to get that particular line of code in MCU (NXP controller), which is executed before Reset
- How to access data from ozone using spark?
- How to create a Hive table using Ozone?
- Pod cannot mount Persistent Volume created by ozone CSI provisioner
- Prometheus Integration with Hadoop (Ozone Cluster)
- Using Apache Ozone FileSystem API results in error
- Apache Ozone + AWS S3 .Net API: PutObject is creating a bucket instead of a key
Trending Questions
- UIImageView Frame Doesn't Reflect Constraints
- Is it possible to use adb commands to click on a view by finding its ID?
- How to create a new web character symbol recognizable by html/javascript?
- Why isn't my CSS3 animation smooth in Google Chrome (but very smooth on other browsers)?
- Heap Gives Page Fault
- Connect ffmpeg to Visual Studio 2008
- Both Object- and ValueAnimator jumps when Duration is set above API LvL 24
- How to avoid default initialization of objects in std::vector?
- second argument of the command line arguments in a format other than char** argv or char* argv[]
- How to improve efficiency of algorithm which generates next lexicographic permutation?
- Navigating to the another actvity app getting crash in android
- How to read the particular message format in android and store in sqlite database?
- Resetting inventory status after order is cancelled
- Efficiently compute powers of X in SSE/AVX
- Insert into an external database using ajax and php : POST 500 (Internal Server Error)
Popular # Hahtags
Popular Questions
- How do I undo the most recent local commits in Git?
- How can I remove a specific item from an array in JavaScript?
- How do I delete a Git branch locally and remotely?
- Find all files containing a specific text (string) on Linux?
- How do I revert a Git repository to a previous commit?
- How do I create an HTML button that acts like a link?
- How do I check out a remote Git branch?
- How do I force "git pull" to overwrite local files?
- How do I list all files of a directory?
- How to check whether a string contains a substring in JavaScript?
- How do I redirect to another webpage?
- How can I iterate over rows in a Pandas DataFrame?
- How do I convert a String to an int in Java?
- Does Python have a string 'contains' substring method?
- How do I check if a string contains a specific word?
tl;dr Yes, Ozone Client (Used by Apache Spark) will prefer reading from local node if the block is present on the same node.
Apache Spark uses Hadoop Filesystem Client (Which will call Ozone Client) to read data from Ozone.
For reads, Apache Ozone will sort the block list based on the distance from the client node (if network topology is configured, the sorting will be done based on the network topology).
If Apache Ozone and Apache Spark are co-located and there is a local copy of the block where Apache Spark is running, Ozone client will prefer reading the local copy. In case if there is no local copy, the read will go over network (if network topology is configured, the Ozone Client will prefer blocks from same rack).
This is implemented in HDDS-1586.