问题：

为简单的hadoop mapreduce作业运行两个mapper和两个reducer

殳凯捷

2023-03-14

我一次就完成了这项工作。我以

$hadoop jar job.jar输入输出

我已经开始了

$ hadoop namenode -format
$ hadoop namenode

$ hadoop datanode

package org.apache.hadoop.examples;

import java.io.IOException;
import java.util.StringTokenizer;
import org.apache.commons.logging.Log;
import org.apache.commons.logging.LogFactory;

import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.io.IntWritable;
import org.rg.apache.hadoop.fs.Path;
import oapache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.Reducer;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
import org.apache.hadoop.util.GenericOptionsParser;

public class WordCount {
private static final Log LOG = LogFactory.getLog(WordCount.class);

  public static class TokenizerMapper
       extends Mapper<Object, Text, Text, IntWritable>{

    private final static IntWritable one = new IntWritable(1);
    private Text word = new Text();

    public void map(Object key, Text value, Context context
                    ) throws IOException, InterruptedException {
      StringTokenizer itr = new StringTokenizer(value.toString());
      while (itr.hasMoreTokens()) {
        word.set(itr.nextToken());
        context.write(word, one);
      }
    }
  }

  public static class IntSumReducer
       extends Reducer<Text,IntWritable,Text,IntWritable> {
    private IntWritable result = new IntWritable();

    public void reduce(Text key, Iterable<IntWritable> values,
                       Context context
                       ) throws IOException, InterruptedException {
      int sum = 0;
      //printKeyAndValues(key, values);

      for (IntWritable val : values) {
        sum += val.get();
      LOG.info("val = " + val.get());
      }
      LOG.info("sum = " + sum + " key = " + key);
      result.set(sum);
      context.write(key, result);
      //System.err.println(String.format("[reduce] word: (%s), count: (%d)", key, result.get()));
    }


  // a little method to print debug output
    private void printKeyAndValues(Text key, Iterable<IntWritable> values)
    {
      StringBuilder sb = new StringBuilder();
      for (IntWritable val : values)
      {
        sb.append(val.get() + ", ");
      }
      System.err.println(String.format("[reduce] key: (%s), value: (%s)", key, sb.toString()));
    }
  }

  public static void main(String[] args) throws Exception {
    Configuration conf = new Configuration();
    String[] otherArgs = new GenericOptionsParser(conf, args).getRemainingArgs();
    if (otherArgs.length != 2) {
      System.err.println("Usage: wordcount <in> <out>");
      System.exit(2);
    }
    Job job = new Job(conf, "word count");
    job.setJarByClass(WordCount.class);
    job.setMapperClass(TokenizerMapper.class);
    job.setCombinerClass(IntSumReducer.class);
    job.setReducerClass(IntSumReducer.class);
    job.setOutputKeyClass(Text.class);
    job.setOutputValueClass(IntWritable.class);
    FileInputFormat.addInputPath(job, new Path(otherArgs[0]));
    FileOutputFormat.setOutputPath(job, new Path(otherArgs[1]));

    System.exit(job.waitForCompletion(true) ? 0 : 1);
}
}

现在你们谁能帮我运行两个映射器和简化器来完成这个单词计数工作吗？

共有1个答案

赫连越

2023-03-14

Gladnick：如果您打算使用默认的TextInputFormat格式，那么在输入文件的数量上至少会有同样多的映射器（或者更多取决于文件的大小）。因此，只需将2个文件放入您的输入目录，这样您就可以运行2个映射器了。（建议使用此解决方案，因为您计划将其作为测试用例运行）。

既然您已经要求了2个减速器，那么您所需要做的就是job.SetNumReduceTasks(2)在您的主要befor submiting the job。

之后，只需准备一个应用程序的jar并在hadoop伪集群中运行它。

            Configuration configuration = new Configuration();
        // create a configuration object that provides access to various
        // configuration parameters
        Job job = new Job(configuration, "Wordcount-Vowels & Consonants");
        // create the job object and set job name as Wordcount-Vowels &
        // Consonants
        job.setJarByClass(WordCount.class);
        // set the main class
        job.setNumReduceTasks(2);
        // set the number of reduce tasks required
        job.setMapperClass(WordCountMapper.class);
        // set the map class for the job
        job.setCombinerClass(WordCountCombiner.class);
        // set the combiner class for the job
        job.setPartitionerClass(VowelConsonantPartitioner.class);
        // set the partitioner class for the job
        job.setReducerClass(WordCountReducer.class);
        // set the reduce class for the job
        job.setOutputKeyClass(Text.class);
        // set the output type of key (the word) expected from the job, Text
        // analogous to String
        job.setOutputValueClass(IntWritable.class);
        // set the output type of value (the count) expected from the job,
        // IntWritable analogous to int
        FileInputFormat.addInputPath(job, new Path(args[0]));
        // set the input directory for fetching the input files
        FileOutputFormat.setOutputPath(job, new Path(args[1]));

类似资料：

如何在Google Dataproc上运行两个并行作业

默认情况下，所有设置都是。如何启用同时运行的多个作业？
写一个简单的cron作业来运行一个Java类

我怎么能写一个cron作业从头开始运行一个java类或写一个cron作业类嵌入Java代码运行？我如何设置计时器运行每一分钟（例如）cron作业？注意：完全初学者与Linux
如何在spring batch中同时运行两个作业

我正在尝试使用Spring Batch和Spring Task Scheduler运行两个作业，而不考虑它们的调度时间。这两个作业（Tasklet）在不同的时间间隔执行不同的作业。以下是springConfig。xml文件：以下是CouPonTougleActivationScheduler和OTPJobScheduler的实现： @enableSched调度公有类CouPonTougleAc
Java：将两个整数的和作为两个的串联打印

问题内容：考虑以下代码：运行此代码时，输出为1711。有人可以告诉我如何获得1711吗？问题答案：该是有直接。是一个等于十进制的八进制常量。在字符串后面加在一起时，它们被串联为字符串，而不是数字。我认为您想要的是：要么这将是更好的结果。
Jmeter-如何将两个采样器作为一个包运行

现在我有两个api方法要测试 POST索引成员删除索引成员问题是indexmember的字段必须是唯一的。因此，当我运行POST时但是当我添加更多线程时= 我在考虑让DELETE作为POST的某种子采样器。因此，POST和DELETE将一起放在一个线程中。任何建议将不胜感激。
限制两个作业不能在Quartz-Scheduler中同时运行

在Quartz-Scheduler中是否可以定义作业的执行约束？谢谢你的回答。

为简单的hadoop mapreduce作业运行两个mapper和两个reducer

共有1个答案

相关问答

相关文章

相关阅读

相关工具

相关文档