Data Deredundancy

RRD Operator Overview

The data with the largest difference can be removed based on the preset proportion.

Table 1 Advanced parameters

Name

Mandatory

Default

Description

sample_ratio

No

0.9

Percentage of reserved data. The value ranges from 0 to 1. For example, 0.9 indicates that 90% of the original data is reserved.

n_clusters

auto

auto

Number of data sample types. The default value is auto, indicating that the total number of types is obtained based on the number of images in the directory. For example, you can specify the number of types to 4.

do_validation

No

True

Indicates whether to validate data. The value can be True or False. True indicates that data is validated before deredundancy. False indicates that data is deduplicated only.

Operator Input Requirements

The following two types of operator input are available:

Output Description