Search in sources :

Example 1 with Pipeline

use of org.talend.dataprep.transformation.pipeline.Pipeline in project data-prep by Talend.

the class ActionTestWorkbench method test.

public static void test(Collection<DataSetRow> input, AnalyzerService analyzerService, ActionRegistry actionRegistry, RunnableAction... actions) {
    final List<RunnableAction> allActions = new ArrayList<>();
    Collections.addAll(allActions, actions);
    final DataSet dataSet = new DataSet();
    final RowMetadata rowMetadata = input.iterator().next().getRowMetadata();
    final DataSetMetadata dataSetMetadata = new DataSetMetadata();
    dataSetMetadata.setRowMetadata(rowMetadata);
    dataSet.setMetadata(dataSetMetadata);
    dataSet.setRecords(input.stream());
    final TestOutputNode outputNode = new TestOutputNode(input);
    Pipeline pipeline = // 
    Pipeline.Builder.builder().withActionRegistry(actionRegistry).withInitialMetadata(rowMetadata, // 
    true).withActions(// 
    allActions).withAnalyzerService(analyzerService).withStatisticsAdapter(// 
    new StatisticsAdapter(40)).withOutput(// 
    () -> outputNode).build();
    pipeline.execute(dataSet);
    // Some tests rely on the metadata changes in the provided metadata so set back modified columns in row metadata
    // (although this should be avoided in tests).
    // TODO Make this method return the modified metadata iso. setting modified columns.
    rowMetadata.setColumns(outputNode.getMetadata().getColumns());
    for (DataSetRow dataSetRow : input) {
        dataSetRow.setRowMetadata(rowMetadata);
    }
}
Also used : StatisticsAdapter(org.talend.dataprep.dataset.StatisticsAdapter) DataSet(org.talend.dataprep.api.dataset.DataSet) RunnableAction(org.talend.dataprep.transformation.actions.common.RunnableAction) RowMetadata(org.talend.dataprep.api.dataset.RowMetadata) DataSetMetadata(org.talend.dataprep.api.dataset.DataSetMetadata) DataSetRow(org.talend.dataprep.api.dataset.row.DataSetRow) Pipeline(org.talend.dataprep.transformation.pipeline.Pipeline)

Example 2 with Pipeline

use of org.talend.dataprep.transformation.pipeline.Pipeline in project data-prep by Talend.

the class PipelineDiffTransformer method buildExecutable.

/**
 * Starts the transformation in preview mode.
 *
 * @param input the dataset content.
 * @param configuration The {@link Configuration configuration} for this transformation.
 */
@Override
public ExecutableTransformer buildExecutable(DataSet input, Configuration configuration) {
    Validate.notNull(input, "Input cannot be null.");
    final PreviewConfiguration previewConfiguration = (PreviewConfiguration) configuration;
    final RowMetadata rowMetadata = input.getMetadata().getRowMetadata();
    final TransformerWriter writer = writerRegistrationService.getWriter(configuration.formatId(), configuration.output(), configuration.getArguments());
    // Build diff pipeline
    final Node diffWriterNode = new DiffWriterNode(writer);
    final String referenceActions = previewConfiguration.getReferenceActions();
    final String previewActions = previewConfiguration.getPreviewActions();
    final Pipeline referencePipeline = buildPipeline(rowMetadata, referenceActions);
    final Pipeline previewPipeline = buildPipeline(rowMetadata, previewActions);
    // Filter source records (extract TDP ids information)
    final List<Long> indexes = previewConfiguration.getIndexes();
    final boolean isIndexLimited = indexes != null && !indexes.isEmpty();
    final Long minIndex = isIndexLimited ? indexes.stream().mapToLong(Long::longValue).min().getAsLong() : 0L;
    final Long maxIndex = isIndexLimited ? indexes.stream().mapToLong(Long::longValue).max().getAsLong() : Long.MAX_VALUE;
    final Predicate<DataSetRow> filter = isWithinWantedIndexes(minIndex, maxIndex);
    // Build diff pipeline
    final Node diffPipeline = // 
    NodeBuilder.filteredSource(filter).dispatchTo(referencePipeline, // 
    previewPipeline).zipTo(// 
    diffWriterNode).build();
    // wrap this transformer into an ExecutableTransformer
    return new ExecutableTransformer() {

        @Override
        public void execute() {
            // Run diff
            try {
                // Print pipeline before execution (for debug purposes).
                diffPipeline.logStatus(LOGGER, "Before execution: {}");
                input.getRecords().forEach(r -> diffPipeline.exec().receive(r, rowMetadata));
                diffPipeline.exec().signal(Signal.END_OF_STREAM);
            } finally {
                // Print pipeline after execution (for debug purposes).
                diffPipeline.logStatus(LOGGER, "After execution: {}");
            }
        }

        @Override
        public void signal(Signal signal) {
            diffPipeline.exec().signal(signal);
        }
    };
}
Also used : DiffWriterNode(org.talend.dataprep.transformation.pipeline.model.DiffWriterNode) BasicNode(org.talend.dataprep.transformation.pipeline.node.BasicNode) Node(org.talend.dataprep.transformation.pipeline.Node) DiffWriterNode(org.talend.dataprep.transformation.pipeline.model.DiffWriterNode) Pipeline(org.talend.dataprep.transformation.pipeline.Pipeline) Signal(org.talend.dataprep.transformation.pipeline.Signal) PreviewConfiguration(org.talend.dataprep.transformation.api.transformer.configuration.PreviewConfiguration) ExecutableTransformer(org.talend.dataprep.transformation.api.transformer.ExecutableTransformer) RowMetadata(org.talend.dataprep.api.dataset.RowMetadata) TransformerWriter(org.talend.dataprep.transformation.api.transformer.TransformerWriter) DataSetRow(org.talend.dataprep.api.dataset.row.DataSetRow)

Example 3 with Pipeline

use of org.talend.dataprep.transformation.pipeline.Pipeline in project data-prep by Talend.

the class PipelineTransformer method buildExecutable.

@Override
public ExecutableTransformer buildExecutable(DataSet input, Configuration configuration) {
    final RowMetadata rowMetadata = input.getMetadata().getRowMetadata();
    // prepare the fallback row metadata
    RowMetadata fallBackRowMetadata = transformationRowMetadataUtils.getMatchingEmptyRowMetadata(rowMetadata);
    final TransformerWriter writer = writerRegistrationService.getWriter(configuration.formatId(), configuration.output(), configuration.getArguments());
    final ConfiguredCacheWriter metadataWriter = new ConfiguredCacheWriter(contentCache, DEFAULT);
    final TransformationMetadataCacheKey metadataKey = cacheKeyGenerator.generateMetadataKey(configuration.getPreparationId(), configuration.stepId(), configuration.getSourceType());
    final PreparationMessage preparation = configuration.getPreparation();
    // function that from a step gives the rowMetadata associated to the previous/parent step
    final Function<Step, RowMetadata> previousStepRowMetadataSupplier = s -> // 
    Optional.ofNullable(s.getParent()).map(// 
    id -> preparationUpdater.get(id)).orElse(null);
    final Pipeline pipeline = // 
    Pipeline.Builder.builder().withAnalyzerService(// 
    analyzerService).withActionRegistry(// 
    actionRegistry).withPreparation(// 
    preparation).withActions(// 
    actionParser.parse(configuration.getActions())).withInitialMetadata(rowMetadata, // 
    configuration.volume() == SMALL).withMonitor(// 
    configuration.getMonitor()).withFilter(// 
    configuration.getFilter()).withLimit(// 
    configuration.getLimit()).withFilterOut(// 
    configuration.getOutFilter()).withOutput(// 
    () -> new WriterNode(writer, metadataWriter, metadataKey, fallBackRowMetadata)).withStatisticsAdapter(// 
    adapter).withStepMetadataSupplier(// 
    previousStepRowMetadataSupplier).withGlobalStatistics(// 
    configuration.isGlobalStatistics()).allowMetadataChange(// 
    configuration.isAllowMetadataChange()).build();
    // wrap this transformer into an executable transformer
    return new ExecutableTransformer() {

        @Override
        public void execute() {
            try {
                LOGGER.debug("Before transformation: {}", pipeline);
                pipeline.execute(input);
            } finally {
                LOGGER.debug("After transformation: {}", pipeline);
            }
            if (preparation != null) {
                final UpdatedStepVisitor visitor = new UpdatedStepVisitor(preparationUpdater);
                pipeline.accept(visitor);
            }
        }

        @Override
        public void signal(Signal signal) {
            pipeline.signal(signal);
        }
    };
}
Also used : WriterNode(org.talend.dataprep.transformation.pipeline.model.WriterNode) WriterRegistrationService(org.talend.dataprep.transformation.format.WriterRegistrationService) SMALL(org.talend.dataprep.transformation.api.transformer.configuration.Configuration.Volume.SMALL) StepMetadataRepository(org.talend.dataprep.transformation.service.StepMetadataRepository) TransformerWriter(org.talend.dataprep.transformation.api.transformer.TransformerWriter) LoggerFactory(org.slf4j.LoggerFactory) Autowired(org.springframework.beans.factory.annotation.Autowired) Configuration(org.talend.dataprep.transformation.api.transformer.configuration.Configuration) Signal(org.talend.dataprep.transformation.pipeline.Signal) DEFAULT(org.talend.dataprep.cache.ContentCache.TimeToLive.DEFAULT) PreparationMessage(org.talend.dataprep.api.preparation.PreparationMessage) Function(java.util.function.Function) AnalyzerService(org.talend.dataprep.quality.AnalyzerService) ActionParser(org.talend.dataprep.transformation.api.action.ActionParser) CacheKeyGenerator(org.talend.dataprep.cache.CacheKeyGenerator) TransformationMetadataCacheKey(org.talend.dataprep.cache.TransformationMetadataCacheKey) DataSet(org.talend.dataprep.api.dataset.DataSet) Logger(org.slf4j.Logger) ActionRegistry(org.talend.dataprep.transformation.pipeline.ActionRegistry) TransformationRowMetadataUtils(org.talend.dataprep.transformation.service.TransformationRowMetadataUtils) Step(org.talend.dataprep.api.preparation.Step) ConfiguredCacheWriter(org.talend.dataprep.transformation.api.transformer.ConfiguredCacheWriter) ContentCache(org.talend.dataprep.cache.ContentCache) ExecutableTransformer(org.talend.dataprep.transformation.api.transformer.ExecutableTransformer) Component(org.springframework.stereotype.Component) StatisticsAdapter(org.talend.dataprep.dataset.StatisticsAdapter) Optional(java.util.Optional) Pipeline(org.talend.dataprep.transformation.pipeline.Pipeline) Transformer(org.talend.dataprep.transformation.api.transformer.Transformer) RowMetadata(org.talend.dataprep.api.dataset.RowMetadata) TransformationMetadataCacheKey(org.talend.dataprep.cache.TransformationMetadataCacheKey) Step(org.talend.dataprep.api.preparation.Step) Pipeline(org.talend.dataprep.transformation.pipeline.Pipeline) Signal(org.talend.dataprep.transformation.pipeline.Signal) WriterNode(org.talend.dataprep.transformation.pipeline.model.WriterNode) ExecutableTransformer(org.talend.dataprep.transformation.api.transformer.ExecutableTransformer) RowMetadata(org.talend.dataprep.api.dataset.RowMetadata) PreparationMessage(org.talend.dataprep.api.preparation.PreparationMessage) TransformerWriter(org.talend.dataprep.transformation.api.transformer.TransformerWriter) ConfiguredCacheWriter(org.talend.dataprep.transformation.api.transformer.ConfiguredCacheWriter)

Aggregations

RowMetadata (org.talend.dataprep.api.dataset.RowMetadata)3 Pipeline (org.talend.dataprep.transformation.pipeline.Pipeline)3 DataSet (org.talend.dataprep.api.dataset.DataSet)2 DataSetRow (org.talend.dataprep.api.dataset.row.DataSetRow)2 StatisticsAdapter (org.talend.dataprep.dataset.StatisticsAdapter)2 ExecutableTransformer (org.talend.dataprep.transformation.api.transformer.ExecutableTransformer)2 TransformerWriter (org.talend.dataprep.transformation.api.transformer.TransformerWriter)2 Signal (org.talend.dataprep.transformation.pipeline.Signal)2 Optional (java.util.Optional)1 Function (java.util.function.Function)1 Logger (org.slf4j.Logger)1 LoggerFactory (org.slf4j.LoggerFactory)1 Autowired (org.springframework.beans.factory.annotation.Autowired)1 Component (org.springframework.stereotype.Component)1 DataSetMetadata (org.talend.dataprep.api.dataset.DataSetMetadata)1 PreparationMessage (org.talend.dataprep.api.preparation.PreparationMessage)1 Step (org.talend.dataprep.api.preparation.Step)1 CacheKeyGenerator (org.talend.dataprep.cache.CacheKeyGenerator)1 ContentCache (org.talend.dataprep.cache.ContentCache)1 DEFAULT (org.talend.dataprep.cache.ContentCache.TimeToLive.DEFAULT)1