Database Reference
In-Depth Information
Chapter 9. Advanced Text Processing with
Spark
In Chapter 3 , Obtaining, Processing, and Preparing Data with Spark , we covered various
topics related to feature extraction and data processing, including the basics of extracting
features from text data. In this chapter, we will introduce more advanced text processing
techniques available in MLlib to work with large-scale text datasets.
In this chapter, we will:
• Work through detailed examples that illustrate data processing, feature extraction,
and the modeling pipeline, as they relate to text data
• Evaluate the similarity between two documents based on the words in the docu-
ments
• Use the extracted text features as inputs for a classification model
• Cover a recent development in natural language processing to model words them-
selves as vectors and illustrate the use of Spark's Word2Vec model to evaluate the
similarity between two words, based on their meaning
Search WWH ::




Custom Search