About Release 7.9.0
This site contains documentation for HPE Ezmeral Data Fabric release 7.9.0, including installation, configuration, administration, and reference content, as well as content for the associated ecosystem components and drivers.
7.9.0 Installation
This section contains information about installing HPE Ezmeral Data Fabric software. It also contains information about how to migrate data and applications from an Apache Hadoop cluster to a HPE Ezmeral Data Fabric cluster.
7.9.0 Data Fabric
HPE Ezmeral Data Fabric is the industry-leading data platform for AI and analytics that solves enterprise business needs.
7.9.0 Administration
This section describes how to manage the nodes and services that make up a cluster.
- Administering Users and Clusters
  Lists topics that help manage a Data Fabric cluster.
- Administering Nodes
  Provides a synopsis of managing nodes in a cluster.
- Administering Volumes
  This section provide information about how to organize and manage data using volumes, a unique feature of HPE Ezmeral Data Fabric clusters.
- Administering Files and Directories
- Administering Tables
  Administration of the HPE Ezmeral Data Fabric Database is done primarily via the command line (maprcli) or with the Managed Control System (MCS). Regardless of whether the HPE Ezmeral Data Fabric Database table is used for binary files or JSON documents, the same types of commands are used with slightly different parameter options. HPE Ezmeral Data Fabric Database administration is associated with tables, columns and column families, and table regions.
- Administering Streams
- Administering Data Fabric Gateways
  A HPE Ezmeral Data Fabric gateway mediates one-way communication between a source HPE Ezmeral Data Fabric cluster and a destination cluster. You can replicate HPE Ezmeral Data Fabric Database tables (binary and JSON) and HPE Ezmeral Data Fabric Streams streams. HPE Ezmeral Data Fabric gateways also apply updates from JSON tables to their secondary indexes and propagate Change Data Capture (CDC) logs.
- Administering Services
- Monitoring the Cluster
  This section describes how to monitor the health and performance of a MapR cluster.
- Configuring Security
  Describes how to configure security and manage secure clusters.
- Managing Secure Clusters
  Provides procedures that will enable you to use Data Fabric clusters securely.
- Administering the Data Access Gateway
  The HPE Ezmeral Data Fabric Data Access Gateway is a service that acts as a proxy and gateway for translating requests between lightweight client applications and the HPE Ezmeral Data Fabric cluster. This section describes considerations when upgrading the service, how to modify configuration settings, and how to administer and manage the service.
- Planning for High Availability
  - CLDB Failover
    Explains the concept of CLDB failover, and its advantages.
  - Best Practices for Running a Highly Available Cluster
    Lists high availability cluster replication types, and the best practices for running such a cluster.
    - Recommended Settings to Recover from Unplanned Shutdown
      - Enabling Fast Failover of Services
        Describes the Fast Failover feature that allows a cluster to rapidly detect and recover from network failures.
      - Tuning the TCP for Fast Failure Detection
        Describes how to tune the TCP stack to detect node or network failures rapidly.
      - Reducing Failure Detection Time for File Clients
        Describes how to set the time for Hadoop and POSIX clients to detect node failures.
    - Recommended Settings for Planned Shutdown
      Explains the modalities of a planned shutdown.
  - ResourceManager High Availability
    Provides an overview of how high availability for Resource Manager works.
- Administrator's Reference
  This section contains in-depth reference information for the administrator.
- Troubleshooting Cluster Administration
  Lists the common errors and their solutions.
- Best Practices for Backing Up HPE Ezmeral Data Fabric Information
  Lists the best practices and performance considerations to follow when backing up HPE Ezmeral Data Fabric information.
- IPv6 Support in Data Fabric
  Describes the IPv6 support feature for Data Fabric.
7.9.0 Development
This section contains information related to application development for Ezmeral ecosystem components and HPE Ezmeral Data Fabric products, including the file system, Database (Key-Value and JSON), and Event Streams.
Other Docs
This section contains release-independent information, including: Installer documentation, Ecosystem release notes, interoperability matrices, security vulnerabilities, and links to other Data Fabric version documentation.
Glossary
Definitions for commonly used terms in MapR Converged Data Platform environments.

Tuning the TCP for Fast Failure Detection

Describes how to tune the TCP stack to detect node or network failures rapidly.

An unplanned failure chiefly takes the form of a node failure or a network failure. In both instances, the network layer retries to connect to the failed node. The number of retry attempts is dictated by the TCP parameter /proc/sys/net/ipv4/tcp_syn_retries. The default value of that parameter is 5 (in Linux), resulting in a latency of more than a minute to detect the node failure. The problem is compounded when the same failed node is contacted repeatedly in the context of a long operation, such as when a client accesses multiple data objects present on that node.

The Data Fabric stack solves the problem by remembering (caching) the information about a node’s failure, and by not contacting that node for subsequent operations on data objects present on that node. Since all form of data is replicated, Data Fabric services find alternative locations for a data object. This feature is in-built into the current software and does not have to be enabled explicitly. Hence, the communication between a client and a recently failed node incurs a one-time long-duration latency. As mentioned before, that latency is governed by the number of retries at the TCP level. Hence, to further improve the one-time longer latency of an operation between a pair of nodes, it is recommended that the number of TCP retries be decreased from 5 to 4, resulting in a latency of about 30 seconds.

Setting the Timeout for TCP Connections

To set the TCP retry count, set the value of tcp_syn_retries to 4 in the /proc/sys/net/ipv4/ directory (for IPv4 connections). For example:

echo 4 > /proc/sys/net/ipv4/tcp_syn_retries

Similarly for IPv6 connections, set:

echo 4 > /proc/sys/net/ipv6/tcp_syn_retries

This TCP setting of 4 ensures that the TCP stack takes about 30 seconds to detect failure of a remote node. To ensure that this setting is persistent across system reboots, set this value in the /etc/sysctl.conf file.

WARNING

This setting impacts all TCP connections to and from a node. Hence, caution must be exercised when lowering this further. Also, in some instances, reducing this further may result in a node being incorrectly flagged as unavailable.

Partners Support Dev-Hub Community ALA Privacy Policy Glossary

HPE Ezmeral Data Fabric – Customer-Managed 7.9.0 Documentation
Abstract	This site contains documentation for the customer-managed platform of the HPE Ezmeral Data Fabric version 7.9.0 including installation, configuration, administration, and reference content, as well as content for the associated bundled ecosystem components and drivers.
Published	April 2025
Edition	7.9.0
Topic last updated	2025-03-07