Deploy Yarn Optimizer
For the YARN Optimizer to function, the Kerberos keytab for the Pulse product must be granted the permission. For more information, see Apache Hadoop 3.3.6 - Yarn Commands.yarn rmadmin -updateNodeResource
The YARN optimizer can be safely deployed by setting the subset to include nodes available in the cluster for optimization.
The behavior of the optimization algorithm can be controlled using the following settings:
- Reserved buffer memory for non-YARN processes
- Maximum percentage of memory to be overcommitted
- Maximum step size for adjusting the amount of overcommitted memory every 5 seconds
To deploy the Yarn Optimizer server, perform the following steps:
Deploying the Add-on
- Execute the command
to start the add-on deployment process.accelo deploy addons
accelo deploy addons
- From the list of available components, choose the Yarn Optimizer service by navigating with the arrow keys and selecting it.
[root@sac04:~ (ad-default)]$ accelo deploy addons
WARN: Gauntlet is running in dry run mode. Disable this to delete indices from elastic and purge data from mongo DB
INFO: Active Cluster: zeus
? Select the components you would like to install: [Use arrows to move, space to select, <right> to all, <left> to none, type to filter]
[ ] Oozie Connector
[ ] Proxy
[ ] QUERY ROUTER DB
[ ] Recommendation Service
[ ] SHARD SERVER DB
[ ] StandAlone Connector
> [X] Yarn optimizer
- Press Enter to deploy the
container.Yarn Optimizer
Generating the Docker Configuration File:
- Generate the Docker configuration file by running
. This action creates a necessary configuration file for the Yarn Optimizer.accelo admin makeconfig ad-yarn-optimizer
accelo admin makeconfig ad-yarn-optimizer
- Once generated, you'll need to review and possibly edit the file located at
. This step ensures the settings align with your specific requirements./data01/acceldata/config/docker/addons/ad-yarn-optimizer.yml
Output:
[root@sac04:~ (ad-default)]$ accelo admin makeconfig ad-yarn-optimizer
WARN: Gauntlet is running in dry run mode. Disable this to delete indices from elastic and purge data from mongo DB
✓ Done, Configuration file generated
IMPORTANT: Please edit/verify the file '/data01/acceldata/config/docker/addons/ad-yarn-optimizer.yml'.
If the addon is already up and running, use './accelo deploy addons' to remove and recreate the addon service.
[root@sac04:~ (ad-default)]$ cat /data01/acceldata/config/docker/addons/ad-yarn-optimizer.yml
version: "1"
services:
ad-yarn-optimizer:
image: ad-yarn-optimizer
container_name: ""
environment:
- MONGO_URI=ZN4v8cuUTXYvdnDJIDp+R8Z+ZsVXXjv8zDOvh8UwQXqyScAm+LrS8Y9EWT8A8/30
- NATS_HOST=ad-events
volumes:
- /etc/localtime:/etc/localtime:ro
- /data01/acceldata/config/krb/security:/krb/security
- /etc/hosts:/etc/hosts:ro
ulimits: {}
ports:
- 19888:9888
depends_on: []
opts: {}
restart: ""
extra_hosts: []
network_alias: []
label: Yarn optimizer
Restarting the Docker Container:
- To restart the Yarn Optimizer container at any point, use the command
.accelo restart ad-yarn-optimizer
accelo restart ad-yarn-optimizer
Deploy the Yarn Metrics Agent via Ansible
Deploying the Yarn Metrics Agent:
- Ensure any older versions of the agents are uninstalled by executing
.accelo uninstall remote
accelo uninstall remote
- Deploy the new agents using
.accelo deploy addons
accelo deploy addons
Configuring the Yarn Metrics Agent:
Info
- For enabling or disabling the Yarn Metrics Agent during deployment, utilize the HYDRA HOSTS feature flag.
- For agents already installed (Pulse Version 3.3.20), the enable/disable feature is managed through the VARS YAML file.
Using the HYDRA HOSTS YAML File:
Info
This file allows for the copying of the default file to predetermined locations. Ansible will decide whether to start the agent based on the enabled or disabled state of the feature flag.
To modify the flag:yarn_opt_enable
- Open the
file located athydra_hosts.yml.$AcceloHome/work/<clusterName>/hydra_hosts.yml - Change the
value under the vars section toyarn_opt_enable.true - Save your changes.
- Uninstall the agents that are already running using
.accelo uninstall remote - Deploy the agents again with the new config:
.accelo deploy hydra
Using the VARS YAML File:
Important
- This file generates service and configuration files from the
data. If the agents ofvars.ymlare already installed with3.3.20flag asyarn_opt_enableinfalse, to start thehydra_hosts.ymlagent, use thepulseyarnmetricsfile to control the same.vars.yml - For the CDP environment, it is mandatory to configure the correct node ID port. for details, see Configure the OverCommit Timeout Value and Node ID Port.
To enable the Yarn Metrics Agent when the flag is set to yarn_opt_enable:false
- Create an
file if it doesn't exist:override.yml
touch $AcceloHome/work/<clustername>/override.yml
- Open the
file and add the following configuration, ensuring theoverride.ymlis set toyarn_opt_enable.true
vi $AcceloHome/work/<clustername>/override.yml
base:
yarn_opt_enable: true
- Save the file and run
to apply the new configuration.accelo reconfig cluster
accelo reconfig cluster
Verifying the Agent's Status:
To confirm if the agent is operational, perform the following steps:
- Log into one of the YARN Node Manager nodes and run
.systemctl status pulseyarnmetrics
systemctl status pulseyarnmetrics
- For log details, use
.journalctl -u pulseyarnmetrics
journalctl -u pulseyarnmetrics
Sample Output:
[root@hdp201 ~]# journalctl -u pulseyarnmetrics
-- Logs begin at Mon 2024-01-15 17:22:26 IST, end at Fri 2024-03-22 15:47:45 IST. --
Mar 21 22:10:23 hdp201.acceldata.dvl systemd[1]: Started YARN metrics collector service.
Mar 21 22:10:23 hdp201.acceldata.dvl yarnmetrics[37475]: Cannot load the yarnMetric config: '/opt/pulse/yarnmetrics/config/yarnmetrics.conf'. Because: unexpecte
Mar 21 22:10:25 hdp201.acceldata.dvl systemd[1]: pulseyarnmetrics.service holdoff time over, scheduling restart.
Mar 21 22:10:25 hdp201.acceldata.dvl systemd[1]: Stopped YARN metrics collector service.
Mar 21 22:10:25 hdp201.acceldata.dvl systemd[1]: Started YARN metrics collector service.
Mar 21 22:10:25 hdp201.acceldata.dvl yarnmetrics[37594]: Cannot load the yarnMetric config: '/opt/pulse/yarnmetrics/config/yarnmetrics.conf'. Because: unexpecte
Mar 21 22:10:26 hdp201.acceldata.dvl systemd[1]: pulseyarnmetrics.service holdoff time over, scheduling restart.
Mar 21 22:10:26 hdp201.acceldata.dvl systemd[1]: Stopped YARN metrics collector service.
Mar 21 22:10:26 hdp201.acceldata.dvl systemd[1]: Started YARN metrics collector service.
Mar 21 22:10:26 hdp201.acceldata.dvl yarnmetrics[37656]: Cannot load the yarnMetric config: '/opt/pulse/yarnmetrics/config/yarnmetrics.conf'. Because: unexpecte
Mar 21 22:10:27 hdp201.acceldata.dvl systemd[1]: pulseyarnmetrics.service holdoff time over, scheduling restart.
Mar 21 22:10:27 hdp201.acceldata.dvl systemd[1]: Stopped YARN metrics collector service.
Mar 21 22:10:27 hdp201.acceldata.dvl systemd[1]: Started YARN metrics collector service.
Mar 21 22:10:27 hdp201.acceldata.dvl yarnmetrics[37759]: Cannot load the yarnMetric config: '/opt/pulse/yarnmetrics/config/yarnmetrics.conf'. Because: unexpecte
Mar 21 22:10:29 hdp201.acceldata.dvl systemd[1]: pulseyarnmetrics.service holdoff time over, scheduling restart.
Mar 21 22:10:29 hdp201.acceldata.dvl systemd[1]: Stopped YARN metrics collector service.
Mar 21 22:10:29 hdp201.acceldata.dvl systemd[1]: Started YARN metrics collector service.
Mar 21 22:10:29 hdp201.acceldata.dvl yarnmetrics[37942]: Cannot load the yarnMetric config: '/opt/pulse/yarnmetrics/config/yarnmetrics.conf'. Because: unexpecte
Mar 21 22:10:30 hdp201.acceldata.dvl systemd[1]: pulseyarnmetrics.service holdoff time over, scheduling restart.
Deploy Agents via HYStaller
If you are installing agents through HYStaller, follow these instructions:
Environment Variable for Service Control: The variable is used to enable or disable the service through HYStaller.YARN_OPT_ENABLE
- Obtain HYStaller:
- Download the HYStaller binary from the License UI.
- Install Hydra Agent:
- Before installation, replace
with the fully qualified domain name (FQDN) of your Pulse server.<PULSE_HOSTNAME> - Execute the script below to stop and disable existing services before proceeding with the HYStaller installation:
- Before installation, replace
((sudo systemctl stop hydra && sudo systemctl disable hydra || true) && (sudo systemctl stop pulsenode && sudo systemctl disable pulsenode || true) && (sudo systemctl stop pulsejmx && sudo systemctl disable pulsejmx || true) && (sudo systemctl stop pulselogs && sudo systemctl disable pulselogs || true) || true)
sudo chown root:root /tmp/hystaller
sudo chmod 0700 /tmp/hystaller
ls -l /tmp/hystaller
sudo /tmp/hystaller uninstall
- Configure and Install: Set up your environment variables and execute the HYStaller with the install command:
PULSE_HOME="/opt/pulse"
PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin"
HYDRA_SERVER_URL="http://<PULSE_HOSTNAME>:19072"
HYDRA_HEARTBEAT_DURATION="60"
HYDRA_PARCEL_MODE="False"
HYDRA_HOSTNAME_CASE="lower"
HYDRA_HOSTNAME_METHOD="CMD"
HYDRA_HEARTBEAT_JITTER="10"
YARN_OPT_ENABLE="false"
sudo env "PULSE_HOME=$PULSE_HOME" "PATH=$PATH" "HYDRA_SERVER_URL=$HYDRA_SERVER_URL" "HYDRA_HEARTBEAT_DURATION=$HYDRA_HEARTBEAT_DURATION" "HYDRA_PARCEL_MODE=$HYDRA_PARCEL_MODE" "HYDRA_HOSTNAME_CASE=$HYDRA_HOSTNAME_CASE" "HYDRA_HOSTNAME_METHOD=$HYDRA_HOSTNAME_METHOD" "HYDRA_HEARTBEAT_JITTER=$HYDRA_HEARTBEAT_JITTER" /tmp/hystaller install
Configure User Account for Running the Pulseyarnmetrics Agent
Info
This is only required if the agent is not running properly and you want to change the user for it.
Perform the following steps:
- Update the User Configuration in
:vars.yml- Locate the
parameter within theagentfile and change thevars.ymlvalue to the desired username (default is "yarn"):yarnmetrics_user
- Locate the
agent:
yarnmetrics_user: newuser
- Create or Update the
File:override.yml
- If absent, generate a new
file within your cluster's working directory:override.yml
touch $AcceloHome/work/<clustername>/override.yml
- Edit the
file to include the updated user configuration:override.yml
vi $AcceloHome/work/<clustername>/override.yml
- Ensure the file contains the following lines, adjusting the
as necessary:yarnmetrics_user
agent:
yarnmetrics_user: newuser
- Apply Configuration Changes: Save the modifications and execute the
command to update the cluster with the new user settings.accelo reconfig cluster
accelo reconfig cluster
- Verify the Change:
- Check the systemd unit file on one of the nodes to confirm the new user is specified:
cat /etc/systemd/system/pulseyarnmetrics.service
- The
field in the service configuration should reflect the new user account chosen.User
Configure the OverCommit Timeout Value and Node ID Port
Follow the steps to configure the overcommit timeout value to revert the changes if NATs are unreachable. Also, you can configure the Node ID port. This configuration helps you to revert each node to its original value if nats is unreachable for a specific amount of time.
To fetch the node ID port run the command in any cluster node.yarn node -list
The parameters in file are as follows. By default, the OverCommit timeout value is set to 300 seconds.vars.yml
base:
yarn_opt_overcommit_timeout_enabled: "true"
yarn_opt_overcommit_timeout_seconds: "300"
yarn_nodeid_port: "45454"
- Create the
file if not present already.override.yml
touch $AcceloHome/work/<clustername>/override.yml
- Open the
file.override.yml
vi $AcceloHome/work/<clustername>/override.yml
- Put the following configuration in the
file if not present already.override.yml
base:
yarn_opt_overcommit_timeout_enabled: "true"
yarn_opt_overcommit_timeout_seconds: "100"
yarn_nodeid_port: "8041"
- Save the file.
- Run the
cluster.reconfig
accelo reconfig cluster
- Verify the config file for the pulseyarnmetrics agent.
cat /opt/pulse/yarnmetrics/config/yarnmetrics.conf
{
"clusterName": "odp_zoro",
"resourcemanagerIP": "odp102.acceldata.dvl",
"resourcemanagerPort": 8088,
"isKerberosEnabled": false,
"keytabPath": "/opt/pulse/node/config/node.keytab",
"kerberosPrinciple": "hdfs@ACCELDATA.COM",
"kerberosDateFormat": "01/02/2006",
"victoriaDBEndpoint": "http://plat02.acceldata.dvl:19043/insert/1385609323/influx",
"natsEndpoint": "http://plat02.acceldata.dvl:19009",
"natsConsumerAckWait": 20,
"isMetricsProcessingNode": false
"overcommit_timeout_enabled": true,
"overcommit_timeout_seconds": 100,
"nodeIdPort": 8041
}
Enable REST API call for Overcommitment of Node Memory
To enable the Rest API call, perform the following steps:
- Run the following command.
accelo admin makeconfig ad-yarn-optimizer
- Add the following environment variable in the ad-yarn-optimizer.yaml file.
IS_REST_EVENT_SUPPORTED=true
Enable Node Limit as Configured in YARN Configurations
You can follow the below steps to consider the memory limit configured in the Ambari UI > Yarn Configurations.
Follow the below steps to enable this feature:
- Set the environment variable to true in
.ad-yarn-optimizer.yml
ENABLE_LIMIT YARN_STATIC_CONFIG to true
- Restart the service using the following command.
accelo restart ad-yarn-optimizer
Enable Static Queue Optimization
Follow the steps to enable the static Queue Optimization.
- SET "
" environment variable to True in ad-yarn-optimizer.yml.ENABLE_STATIC_QUEUES_OVERCOMMITMENT - In case of the CDP setup, set "
" to "QUEUE_UPDATE_API" in ad-yarn-optimizer.yml./ws/v1/cluster/scheduler-conf?user.name=hdfs - Restart the
service by running the following command.ad-yarn-optimizer
accelo restart ad-yarn-optimizer
- Set
to True in the FEATURE_FLAGS environment variable in ad-graphql.enableStaticQueue

Have a suggestion?