Adding a Ranger Access Permission Policy for Spark
Scenarios
Ranger administrators can use Ranger to set permissions for Spark users, implementing fine-grained permission control over Hive data tables to ensure service security and availability.
Notes and Constraints
- After enabling or disabling Ranger authentication for Spark or Spark2x, you must restart the Spark or Spark2x service for the changes to take effect.
- After modifying Ranger authentication, you must download the client again or manually update the spark-defaults.conf configuration file in the Client installation directory/Spark2x/spark/conf directory.
- To enable Ranger authentication, set spark.ranger.plugin.authorization.enable to true and change the value of spark.sql.authorization.enabled to true.
- To disable Ranger authentication, set spark.ranger.plugin.authorization.enable to false.
- In Spark and Spark2x, spark-beeline (used to connect to JDBCServer) supports Ranger IP address filtering (configured via Policy Conditions in Ranger), but this feature is not supported in spark-submit and spark-sql.
- In MRS 3.3.0-LTS and later versions, the Spark2x service is renamed Spark, and the names of its service roles have also changed. For example, JobHistory2x is renamed JobHistory. The specific service and role names used in descriptions and operations depend on your actual cluster version.
- After adding a Spark SQL access policy in Ranger, you must add the corresponding path access policies in HDFS. Otherwise, data files cannot be accessed. For details, see Adding a Ranger Access Permission Policy for HDFS.
- The global policy in Ranger is used exclusively to associate with the Temporary UDF Admin permission to control UDF package uploads.
- When Ranger is used to control Spark SQL permissions, the empower syntax is not supported.
- Ranger policies do not support local paths or HDFS paths that contain spaces.
- When Ranger authentication is enabled, you must have default permissions for the underlying tables to perform operations on a view. To enable independent authentication for a view without checking table permissions, set the spark.ranger.plugin.viewaccesscontrol.enable parameter to true.
- When submitting jobs using a tool other than spark-beeline, you must set this parameter in the Client installation directory/Spark/spark/conf/spark-defaults.conf file.
- When submitting jobs using spark-beeline, you must set this parameter in the Client installation directory/Spark/spark/conf/spark-defaults.conf file. You must also add this parameter on MRS Manager by choosing JDBCServer > Customization.
Prerequisites
- The Ranger service has been installed and is running properly in the cluster.
- The users, user groups, or roles requiring permissions must already be created in the cluster. For newly created users, you can configure permission policies only after the users are automatically synchronized to Ranger.
- The related user has been added to the hive user group.
- The Ranger authentication function has been enabled for the Hive service in the cluster. If you need to manually enable this function, perform the following operations in sequence: enable Ranger authentication for Hive, restart Hive, restart Spark, enable Ranger authentication for Spark, and then restart Spark again.
Procedure
- Log in to the Ranger web UI as the Ranger administrator. For details, see Logging In to the Ranger Web UI.
- On the home page, click the component plug-in name in the HADOOP SQL area, for example, Hive.

- On the Access tab page, click Add New Policy to add a Spark permission control policy.

- Configure the parameters listed in the table below based on the service demands.
Table 1 Spark permission parameters Parameter
Description
Policy Name
Policy name, which can be customized and must be unique in the service.
Policy Conditions
IP address filtering policy, which can be customized. You can enter one or more IP addresses or IP address segments. An IP address can contain the wildcard character (*), for example, 192.168.1.10,192.168.1.20 or 192.168.1.*.
Policy Label
A label specified for the current policy. You can search for reports and filter policies based on labels.
database
Name of the Spark database to which the policy applies.
The Include policy applies to the current input object, and the Exclude policy applies to objects other than the current input object.
table
Name of the Spark table to which the policy applies.
To add a UDF-based policy, switch to UDF and enter the UDF name.
The Include policy applies to the current input object, and the Exclude policy applies to objects other than the current input object.
column
Name of the column to which the policy applies. The value * indicates all columns.
The Include policy applies to the current input object, and the Exclude policy applies to objects other than the current input object.
Description
Policy description.
Audit Logging
Whether to generate an audit log when a request matches the policy.
- Yes: An audit log is generated whenever an access request matches the policy, regardless of whether the result is Allow or Deny.
- No: No audit log is generated when an access request matches the policy.
Allow Conditions
Policy allow conditions, which define the permissions and exceptions authorized by this policy.
In the Select Role, Select Group, and Select User columns, choose the specific roles, user groups, or users to whom the permissions will be granted. Click Add Conditions and specify the IP address range to which this policy applies. Click Add Permissions to assign the corresponding access rights.
- select: permission to query data
- update: permission to update data
- Create: permission to create data
- Drop: permission to drop data
- Alter: permission to alter data
- Index: permission to index data
- All: all permissions
- Read: permission to read data
- Write: permission to write data
- Temporary UDF Admin: temporary UDF management permission
- Select/Deselect All: permission to select or deselect all
To add multiple permission control rules, click
.If users or user groups in the current condition need to manage this policy, select Delegate Admin. These users will become the agent administrators. The agent administrators can update and delete this policy and create sub-policies based on the original policy.
Deny Conditions
Policy deny conditions, which define the permissions and exceptions that must be rejected by the policy. The configuration method is identical to that of Allow Conditions.
Table 2 Common permission configuration scenarios Task
Role Authorization
role admin operation
- On the home page, click Settings and choose Roles > Add New Role.
- Set Role Name to admin. In the Users area, click Select User and select a username.
- Click Add Users, select Is Role Admin in the row where the username is located, and click Save.
After assigning the Hive administrator role to a user, perform the following operations during each maintenance session:
- Log in to the node where the cluster client is installed as the client installation user.
- Run the following command to configure environment variables:
source Client installation directory/bigdata_env
- Run the following commands to authenticate the user:
kinit Spark service user
- Run the following command to log in to the client tool:
spark-beeline
- Run the following command to update the administrator permissions:
set role admin;
Creating a database table
- Enter the policy name in Policy Name.
- Enter and select the corresponding database on the right of database. (If you want to create a database, enter the name of the database to be created or enter * to indicate a database with any name, and then select the name.) Enter and select the corresponding table name on the right of table and column. Wildcard characters (*) are supported.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Click Add Permissions and select Create.
Deleting a table
- Enter the policy name in Policy Name.
- Enter and select the corresponding database on the right of database. (If you want to delete a database, enter the name of the database to be created or enter * to indicate a database with any name, and then select the name.) Enter and select the corresponding table name on the right of table and column. Wildcard characters (*) are supported.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Click Add Permissions and select Drop.
ALTER operation
- Enter the policy name in Policy Name.
- Enter and select the corresponding database on the right of database, enter and select the corresponding table on the right of table, and enter and select the corresponding column name on the right of column. Wildcard characters (*) are supported.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Click Add Permissions and select Alter.
LOAD operation
- Enter the policy name in Policy Name.
- Enter and select the corresponding database on the right of database, enter and select the corresponding table on the right of table, and enter and select the corresponding column name on the right of column. Wildcard characters (*) are supported.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Click Add Permissions and select update.
INSERT operation
- Enter the policy name in Policy Name.
- Enter and select the corresponding database on the right of database, enter and select the corresponding table on the right of table, and enter and select the corresponding column name on the right of column. Wildcard characters (*) are supported.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Click Add Permissions and select update.
- The user also needs to have the submit-app permission of the Yarn task queue. By default, the Hadoop user group has the submit-app permission of all Yarn task queues. For details about how to load a network instance to a cloud connection, see Adding a Ranger Access Permission Policy for YARN.
GRANT operation
- Enter the policy name in Policy Name.
- Enter and select the corresponding database on the right of database, enter and select the corresponding table on the right of table, and enter and select the corresponding column name on the right of column. Wildcard characters (*) are supported.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Select Delegate Admin.
ADD JAR operation
- Enter the policy name in Policy Name.
- Click database, and select global from the drop-down list. On the right of global, enter related information and select *.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Click Add Permissions and select Temporary UDF Admin.
VIEW and INDEX permissions
- Enter the policy name in Policy Name.
- On the right side of database, enter the database name and select the corresponding database. (If you want to delete a database, enter the database name and select *.) On the right side of table, enter a table name and select the view and index names. On the right side of column, enter a Hive column name, and select *.
- In the Allow Conditions area, select a user from the Select User drop-down list.
- Click Add Permissions and select permissions for the user as required.
Operations on other user database tables
- Perform the preceding operations to add the corresponding permissions.
- Grant the read, write, and execution permissions on the HDFS paths of other user database tables to the current user. For details, see Adding a Ranger Access Permission Policy for HDFS.
After Spark SQL access policy is added on Ranger, you need to add the corresponding path access policies in the HDFS access policy. Otherwise, data files cannot be accessed. For details, see Adding a Ranger Access Permission Policy for HDFS.
- The global policy in Ranger is used exclusively to associate with the Temporary UDF Admin permission to control UDF package uploads.
- When Ranger is used to control Spark SQL permissions, the empower syntax is not supported.
- Ranger policies do not support local paths or HDFS paths that contain spaces.
- When Ranger authentication is enabled, you must have default permissions for the underlying tables to perform operations on a view. To enable independent authentication for a view without checking table permissions, set the spark.ranger.plugin.viewaccesscontrol.enable parameter to true.
- When submitting jobs using a tool other than spark-beeline, you must set this parameter in the Client installation directory/Spark/spark/conf/spark-defaults.conf file.
- When submitting jobs using spark-beeline, you must set this parameter in the Client installation directory/Spark/spark/conf/spark-defaults.conf file. You must also add this parameter on MRS Manager by choosing JDBCServer > Customization.
- Click Add to view basic information about the policy in the policy list. After the policy takes effect, check whether the related permissions are normal.
To disable a policy, click
to edit the policy and set the policy to Disabled.If a policy is no longer used, click
to delete it.
Configuring Data Masking for Spark Tables
Ranger supports data masking for Spark tables. It masks sensitive information by processing the results returned by your SELECT operations.
- Set the value of spark.ranger.plugin.masking.enable to true on both the server and the client.
- Server: Log in to FusionInsight Manager, and choose Clusters > Services > Spark or Spark2x. On the displayed page, click Configurations and then All Configurations. Search for spark.ranger.plugin.masking.enable and set its value to true. Save the change and restart the service.
- Client: Log in to the Spark client node, find the {Client installation directory}/Spark/spark/conf/spark-defaults.conf file, and change the value of spark.ranger.plugin.masking.enable to true.
- Log in to the Ranger web UI. Click the component plug-in name, for example, Hive, in the HADOOP SQL area on the homepage.
- On the Masking tab page, click Add New Policy to add a Spark permission control policy.
- Configure the parameters listed in the table below based on the service demands.
Table 3 Spark2x data masking parameters Parameter
Description
Policy Name
Policy name, which can be customized and must be unique in the service.
Policy Conditions
IP address filtering policy, which can be customized. You can enter one or more IP addresses or IP address segments. An IP address can contain the wildcard character (*), for example, 192.168.1.10,192.168.1.20 or 192.168.1.*.
Policy Label
A label specified for the current policy. You can search for reports and filter policies based on labels.
Hive Database
Name of the Spark database to which the current policy applies. Note that you can specify only one database name, and wildcard characters (*) are not supported.
Hive Table
Name of the Spark table to which the current policy applies. Note that you can specify only one table name, and wildcard characters (*) are not supported.
Hive Column
Name of the Spark column to which the current policy applies. Note that you can specify only one column name, and wildcard characters (*) are not supported.
Description
Policy description.
Audit Logging
Whether to generate an audit log when a request matches the policy.
- Yes: An audit log is generated whenever an access request matches the policy, regardless of whether the result is Allow or Deny.
- No: No audit log is generated when an access request matches the policy.
Transfer Mask
Whether the policy is automatically transferred when dynamic masking is enabled.
For details about how to configure Spark dynamic masking, see Configuring Spark Dynamic Masking.
Mask Conditions
In the Select Group and Select User columns, select the user group or user to receive the permissions, click Add Conditions to specify the target IP address range, and then click Add Permissions and select SELECT.
Click Select Masking Option and select a data masking policy.
- Redact: Use x to mask all letters and 0 to mask all digits.
- Partial mask: show last 4: Only the last four characters are displayed. Other characters are masked with x. Masking of Chinese characters is not supported.
- Partial mask: show first 4: Only the first four characters are displayed. Other characters are masked with x. Masking of Chinese characters is not supported.
- Hash: Perform hash calculation for data.
- Nullify: Replace the original value with the NULL value.
- Unmasked(retain original value): The original data is displayed.
- Date: show only year: Only the year information is displayed.
- Custom: You can use any valid Hive UDF (returns the same data type as the data type in the masked column) to customize the policy.
To add a multi-column masking policy, click
.Deny Conditions
Policy deny conditions, which define the permissions and exceptions that must be rejected by the policy. The configuration method is identical to that of Allow Conditions.
- Click Add to view basic information about the policy in the policy list.
Spark Row-Level Data Filtering
Ranger supports row-level filtering for Spark tables when executing SELECT operations.
- Set spark.ranger.plugin.rowfilter.enable to true on both the server and the client.
- Server: Log in to FusionInsight Manager, and choose Clusters > Services > Spark or Spark2x. On the displayed page, click Configurations and then All Configurations. Search for spark.ranger.plugin.rowfilter.enable and set its value to true for all displayed search results. Save the change and restart the service.
- Client: Log in to the Spark client node, find the {Client installation directory}/Spark/spark/conf/spark-defaults.conf file, and change the value of spark.ranger.plugin.rowfilter.enable to true.
- Log in to the Ranger web UI. Click the component plug-in name, for example, Hive, in the HADOOP SQL area on the homepage.
- On the Row Level Filter tab page, click Add New Policy to add a row data filtering policy.
- Configure the parameters listed in the table below based on the service demands.
Table 4 Parameters for filtering Spark row data Parameter
Description
Policy Name
Policy name, which can be customized and must be unique in the service.
Policy Conditions
IP address filtering policy, which can be customized. You can enter one or more IP addresses or IP address segments. An IP address can contain the wildcard character (*), for example, 192.168.1.10,192.168.1.20 or 192.168.1.*.
Policy Label
A label specified for the current policy. You can search for reports and filter policies based on labels.
Hive Database
Name of the Spark database to which the current policy applies. Note that you can specify only one database name, and wildcard characters (*) are not supported.
Hive Table
Name of the Spark table to which the current policy applies. Note that you can specify only one table name, and wildcard characters (*) are not supported.
Description
Policy description.
Audit Logging
Whether to generate an audit log when a request matches the policy.
- Yes: An audit log is generated whenever an access request matches the policy, regardless of whether the result is Allow or Deny.
- No: No audit log is generated when an access request matches the policy.
Row Filter Conditions
In the Select Role, Select Group, and Select User columns, choose the specific roles, groups, or users to whom the permissions will be granted. Click Add Conditions and specify the IP address range to which this policy applies. Click Add Permissions and select Select.
Click Row Level Filter and enter data filtering rules.
For example, to filter out data where the name column equals zhangsan in table A, configure the filtering rule as name <> 'zhangsan'. For more information, see the Apache Ranger official documentation.
To add more rules, click
. - Click Add to view basic information about the policy in the policy list.
- When you use the Spark client to perform a SELECT operation on a table configured with a data masking policy, the system processes the data and displays the masked results.
Helpful Links
- For details about the priorities of different Ranger policies, see Condition Priorities of the Ranger Permission Policy.
- To view information about permission objects in Ranger, such as users, user groups, and roles, see Viewing Ranger User Permission Synchronization Information.
- You can create new users for an MRS cluster in MRS Manager. For details, see Creating an MRS Cluster User.
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot