CIS 493/593 Big Data (3-0-3) |
||
| Course Content |
|
|
|
IMPORTANT NOTE !!! If You Have a 404 Error in any URL in the Class Lecture Notes under http://eecs.csuohio.edu/~sschung/, You have to replace "http" with "https" to be able to access ! 10. April 25, 2023 The Final Exam on Monday May 8 at 4:00PM - 6:00PM In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final //a href="CIS593S23_FinalExamLinkOnly.html"> See the Links for Final Exam Lecture Note Only (Scroll Down to the Later Half of Lecture Note Sections) Topics to Focus on for the Final of CIS493/593 Design of Intelligent System with data pipelining of Big data processing Unstructured Text Processing Techniques, Data Preprocessing Methods Inverted Index TF-IDF, Cosine Similarity NLP Text Analysis methods - POS Tagging Data Preprocessing Methods for Classification with ML Algorithms Classification with Machine Learning Algorithms: Decision Tree and Neural Network Map Reduce Process on Hadoop Distributed File System One Page Note is Allowed to the Final 10. Jan 16, 2023 Information on Midterm Midterm will Be on March 8 at 4:00PM - 5:30PM ! //a href="CIS593Sp23_MidtermOnly.html"> See Lecture Links Only for MidTerm (Scroll down to the Lecture Note Section to See Active links - All the Rest of Links are Disabled) The Topics to Focus On for the Midterm: Basically whichever Lecture notes (Slides) covered in class! Characteristics of Big Data and their Contents, Main Differences between Big Data and Traditional Data, Common Big Data Applications/Architecture, Universal Data Exchange Formats in Object Exchange Model(OEM): Three Common Big Data formats in Semi-Structured Model: - HTML, XML, JSON - Data Model, Syntax of Each Encoding Format, all the related data processing techniques -- DOM, XPath. Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model for Object Relation Mapping (ORM), Conversion between Relation(CSV), XML, and JSON Unstructured Data Processing, Inverted Index, How to Build Inverted Index, Text Cleaning Preprocessing, Basic Natural Language Processing (NLP) Methods in Data Pipeling, POS Tagging Comparison Between Semistructured Database Server and Relational Database Server for Database Management Strategies. Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining NOTE that ONE Page Note is NOT allowed to the Midterm !! Exam Format: 6-7 Main Questions with 2-3 subquestions for a Small/Short Answer Types for Probelm Solving 8. Jan 16, 2023 Only the registered students can access the course blackboard. If you have a problem with your blackboard, please contact the registrar or e-learning center CSU Tech Support to resolve the issue ! Faculty does not control your registration and the course blackboard access in the CSU systems. //a href="https://www.csuohio.edu/center-for-elearning/technical-support"> e-learning CSU Tech Support TA info: 9. Jan 16, 2023: TA Information TA: Yixi Luo Email: luoyixi.cn@gmail.com Office Hours: Mon, Wed 9:50 AM - 11:50 PM Tentatively (Send email to TA ahead to set up a time slot to let him/her know that you are coming) Location: Big Data Analytics Lab: FH 305 or Recommend a ZOOM Meeting for Your Safety ! ZOOM Meeting ID: 306 371 3308 Password: VFw6s8 //a href="https://zoom.us/j/3063713308?pwd=Zko5SkZKZUxwK2dIYWU1QklUc20xdz09"> ZOOM Meeting Link If you have questions in Labs or grading your Labs, Send an email to TA to See During the TA's Office Hours or Schedule a Zoom meeting Dr. Chung's Office Hours: Mon and Wed 1:30PM - 3:30PM Zoom Meeting Only for Safety ! Send Me Email to Set Up a Zoom Meeting. Meeting ID: 859 3867 6332 Email: s.chung@csuohio.edu 6. Jan 16, 2023: Lab Submission: The Output of each lab is your report in Doc file that shows your screen captures of each of your executions with Your Outputs in your web browser or the server returned the correct results in the webpage. your Report in Doc file also should expain all the platform set up, the execution steps, and copy of each source code files ) Each of your screen capture must show your results returned in your web browser of your system to prove that your lab is done correctly !! Submit your Zip file that includes the followings on Blackboard for a timestamp and as a proof 1) Your Lab Report in .doc file that explains all the platform set up procedures, the execution steps, each intermediate output, final outputs, and a copy of each source code files, 2) All your Source files, and output files 5. Jan 16, 2023: See Updated Instructions for File Permission Error 403 in the Lab Section! 4. Jan 16, 2023: Instructions to Set Up Webpage is Posted in the Lab Section! 3. Jan 16, 2023: Each Class Attendance is Required for CIS492/593 ! The ZOOM Meeting Log for Each Class Shows Your Class Attedance There Will Be Random Quizzes or Sign Up Sheet to Check the Attendance As Well ! 2. Jan 16, 2023: You can use any SQL Server for your Database such as MySql or MS SQL Server. See MySql Set up Guide in the Lab section. 1. Jan 16, 2023: The class webpage for this semester will be announced on blackboard: //a href="https://eecs.csuohio.edu/~sschung/CIS593/CIS593Sp23.html"> CIS492/593 Big Data Class Webpage Or You can Reach from Teaching Section of My Website for Big data Research Lab at: //a href="https://eecs.csuohio.edu/~sschung/"> Big Data Research Lab If you have a trouble to display the webpage correctly with MS Internet Explorer, open it with Google Chrome Check the Last Day to Add and Drop here ! //a href="https://www.csuohio.edu/registrar/academic-calendar"> University's Official Academic Calendar for the Semester and the Final Exam schedules Final Exam Schedule: Mon May 10, 4:00 - 6:00 PM Information on Midterm and Final Exams 7. Midterm will Be Tentatively on March 8 ! The Topics to Focus On for the Midterm: Basically whichever Lecture notes (Slides) covered in class! Characteristics of Big Data and their Contents, Main Differences between Big data and Traditional Data, Common Big Data Applications/ architecture, Universal Data Exchange Formats in Object Exchange Model(OEM): Three common big data formats as semi structured data: - HTML, XML, JSON - Data Model, Syntax of Each Encoding Format, all the related processing techniques -- DOM, XPath. Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model Comparison between Relational Data model and Semi-Structured Data Model. Conversion between Relation(CSV), XML, and JSON Unstructured Data Processing, Inverted Index, Text Cleaning, Basic NLP Processing Methods in Data Pipeling, POS Tagging Comparison Between Semistructured Database Server and Relational Database Server for Database Management Strategies. Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining NOTE that ONE Page Note is NOT allowed to the Midterm !! //a href="CIS593S21_FinalExamLinkOnly.html"> See the Links for Final Exam Lecture Note Only Topics to Focus on for the Final of CIS493/593 Design of Intelligent System with data pipelining of Big data processing Unstructured Text Processing Techniques, Data Preprocessing Methods Inverted Index TF-IDF, Cosine Similarity NLP Text Analysis methods - POS Tagging Data Preprocessing Methods for Classification with ML Algorithms Classification with Machine Learning Algorithms: Decision Tree and Neural Network Map Reduce Process on Hadoop Distributed File System One Page Hand Wrtten Note (Each Side) is Allowed to the Final ! Printed Copy of Lecture Notes Are NOT Permitted as One Page Note |
|
Projects on Big Data Processing, Building an Intelligent Web Application with Big Data Analytics The Scope of Group Project Requirements Has Been Adjusted to: (This Adjustment May not Apply to this semester. Will Be Announced) 1. You Can Do One Person Group Project in a Small Scale. 2. a Project with a Small Scale with Some Extension of One of Lab4 - Lab4_3 3. For Those Who Have Aleady Taken CIS660, it is Required to Complete a Full Project with Bid data set and Presentation 4. Implementing Your Final Project with Clientside and Serverside of a Web Application with a Web Based User Interface Is NOT Required. It will Be Counted as Extra Credit. Final Group Project Specification and Instructions: Project Submission Instructions: Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file either on the google drive or One Drive and send the link. (Send me email for access permission for this link !) Submit a Zip file that includes: All of your presentation slides (both in .ppt and .pdf) and Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files on Blackboard by the end of Friday of your presentation week. Include the Problems/Error Encountered and Your Resolutions in Your Report One Submission Per Group Required. Submit a Zip File that Includes All the required Source Files, Input, Output Files, and Final Report (in Doc file) and Presentation Slides (in pptx). Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. Important Notes for Final Project: - 1 - 3 Person Group Project Are Allowed - One Person Final Project is Allowed. You Can Work Alone. - Three Person Group is Allowed. Make Sure to Make the Project in a Bigger Scale - You can Change Your Project Plan/Proposal or the Details even after Your Proposal is submitted until the Deadline of the Status Report. - If You Need to Find Group Members, Use Blackboard Email to Send to the Class, Some will respond to you if they are looking for a group member. The IMDB Data links in the Project List are gone. Check here for Movie Review Data Set. //a href="https://www.kaggle.com/rounakbanik/the-movies-dataset#movies_metadata.csv"> Movie Review Data //a href="https://developers.google.com/youtube/reporting/"> Youtube Analytic Site For Those Who Have Already Taken CIS660, The Final Group Project Should Include Fully Analytic Processing Suggested Projects: For Text Analytics like Sentiment Analysis or Opinion Analysis: NLP Techniques - POS, NER Tagging, Bi-Gram Handling Are Required for Preprocessing. For Document Categorization: by Constructing TF-IDF Vectorization. Inverted Index Building Will Be Plus but Optional. Building Word2Vec Embeddings for a Collection of Documents/Webpages with Training Set Generation in Skip Gram Model For Other Types of Projects, the Proposal is Required to Be Approved to Meet the Complexity of Final Project Extra Credit Final Project: Final Group Project Time Line: (Tentative) Group Project Proposal Due by April 7 ! Group Project Status Report Due by April 21 ! Group Project Presentation Either on May 1 or May 3 ! Final Group Project Report Due By Friday May 5 ! Any Group Size in 1 - 4 Person Is Allowed Since the Size of the Class is Too Big, 3-4 Person Project Group Will Be Allowed As Long As the Project Scale is Big Enough for 3-4 Persons IMPORTANT Submissions for Group Project Task 1: Group Project Proposal (Plan) Submit minimum 3 Page Group Proposal on Blackboard with List of Group Memebers Group Project Proposal (Plan) Should Include Brief Descriptions on: 1. Description of Big Data in Size and Format, Data Collection Plan, 2. Goal of Your Big Data Analytic Project with What Kind of Intelligent Analytic Funtionality, Features of Your AI/Big Data Analytic Application 3. Big Data Processing Plan, Methods 4. Investigate on Platform/Systems/Tools/APIs to Use Task 2: Group Project Status Report Your Project Status Report Should Show the Following Tasks Done: 1. Platform Setting/System Configuration Procedure (if it is new) 2.Your KnowledgeBase Structure/ Database Design 3. Design of Big Data Processing Pipeline, Data Transformation Methods 4. Your Server Data/Database Contents in Progress, Any Intermediate Outputs in Progress Task 3: Project Presentation Project Presentation Starts from the Last Week of the Class of the Semester Project Presentation Schedule Will Be Sent To Your CSU Email for Sign Up One week Before the Presentation Read the Instructions of Project Presentation Here ! Group Project Presentation Should Include: 1. Data Description, Data Size, Data Collection Method 2. Platform Setting/System Configuration Procedures 3. System Design (Architecture) of Your AI Application or Data Analytic Goal in Detail 4. Raw Big Data Preprocessing Methods and Intermediate Results 5. Design of Big Data Processing Pipeline, Data Transformation Methods 6. Description of Your KnowledgeBase Structure/Database Design. And Show the Contents 7. Ranking Algorithm, Data Matrix (Structures) if any for Evaluation 8. The Problems/Errors Encountered and Your Resolutions 9. System Demo or Evaluation Results and Visualization of the Result Submit Final Project Report and Presentation By the End of Friday of The Last Class Week After Your Presentation! One Submission Per Group Required. Submit a Zip file that includes: 1. All of your presentation slides (both in .pptx) and 2. Your Group Final Project Report (in doc) The Final Project Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. The Report Should Include Platform/System Set up the Set Up Procedure /Configuration Detail of Your Platform/System/Packages, Executions Steps, all the Source Codes, Scripts, all the intermediate outputs, and final output files. Include the Problems/Error Encountered and Your Resolutions in Your Report If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. Group Project Presentation Should Include: 1. Data Description, Data Size, Data Collection Method 2. Platform Setting/System Configuration Procedures 3. System Design (Architecture) of Your AI Application or Data Analytic Goal in Detail 4. Raw Big Data Preprocessing Methods and Intermediate Results 5. Design of Big Data Processing Pipeline, Data Transformation Methods 6. Description of Your KnowledgeBase Structure/Database Design. And Show the Contents 7. Ranking Algorithm, Data Matrix (Structures) if any for Evaluation 8. The Problems/Errors Encountered and Your Resolutions 9. System Demo or Evaluation Results and Visualization of the Result < Requirements of Combining two Group projects of two related courses: Combining two Group projects of two related courses (CIS492 with CIS408 or CIS593 Deep Learning) are ok as long as the combined project focuses on both sides of the subjects for example, with Web application aspect for CIS408 and Server-side big data processing with a database server and any analytic methods with big data covered in CIS492. The data of the Deep Learning class CIS593 is images, which is a lot different from the Big data covered in this Big data class such as a large volumn of semistructured collection or text documents. Maybe training a large text documents using deep learning would be a good combined project of CIS492/CIS593 and the Deep Learning Class. For Extra Credit Projects or Contract Course Requirement for Honor Students It includes a project to build a real-life web based AI applications like: - Web Search Engine for a certain domain, for example, .csuohio.edu or .cnn - WebMD - Question Answering system like Alexa but in a webbased interface to get a user question Project Examples of Vectorization of each document in a Training set for a Machine Learning Classifier: //a href="https://eecs.csuohio.edu/~sschung/CIS660/ProjectExample_SentimentAnalysis_YelpReview.pdf"> Project Example I: Sentiment Analysis of Yelp Business Review with Machine Learning //a href="https://eecs.csuohio.edu/~sschung/CIS660/ProjectExample_SentimentAnalysis_YelpReview_Feng.pdf"> Project Example II: Sentiment Analysis of Yelp Business Review with Machine Learning Best Senior Design Projects on Big Data Processing and Text Anaytics (Created from the Subject of CIS493 and CIS408): Best Group Projects: Selected Best Projects Will be Posted Here ! Big Data and Data Science Projects: You Can Choose to Extend One of The Extra Credit Labs Below as a Final Group Project either with a Different data set or the same data set Extra Credit Lab 4_2 on Webpage Categorization by Topics Extra Credit Lab 4_3 on Information Retrieval Methods for Content Based Document Search Engine //a href="https://eecs.csuohio.edu/~sschung/CIS593/Text_Mining_ConstructingInvertedIndex.pdf"> Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example) Extra Credt Lab 5 on Sentiment Analysis with Machine Learning for Classification: Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set Classification Goal: 1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review. 2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star) More To Come Here ! Group Project Data Sources: You can choose to work on these data sets for your group project Final Project Submission Instructions: Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file on your google drive or One Drive and Send email to me and TA to share ! Submit a Zip file on Blackboard by the end of Friday of your presentation week. One Submission Per Group Required. Your Project Zip File Should includes: 1) All of your presentation slides (both in .ppt and .pdf) and 2) Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files 3)Include the Problems/Error Encountered and Your Resolutions in Your Report Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. IMPORTANT NOTE !!! If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. |
|
Basic Python Tutorials : Popular Python Data Science Platforms: //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) • Anaconda //a href="https://www.anaconda.com/open-source"> Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms //a href="https://docs.anaconda.com/anaconda/navigator/tutorials/"> Anaconda Tutorials • Python Anaconda Tutorial Sites //a href="https://data-flair.training/blogs/python-anaconda-tutorial/"> Anaconda Tutorial Site //a href="https://www.edureka.co/blog/python-anaconda-tutorial/"> Anaconda Tutorial Site //a href="https://www.dataquest.io/blog/jupyter-notebook-tutorial/"> Basic Guide for Python tools for Data Science: jupyter-notebook //a href="https://www.dataquest.io/blog/advanced-jupyter-notebooks-tutorial/"> More Basic Guide for Python tools for Data Science: jupyter-notebook • Python Scikit Learn for Common Data Science Tasks Text Preprocessing (Natural Language Processing) Library in Python SpaCy: //a href="https://spacy.io/api"> Liquistic Modules in Python SpaCy //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing • Python sklearn.cluster //a href="https://scikit-learn.org/stable/modules/clustering.html"> Python Sklearn Clustering Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Basic R Tutorials : //a href="IndependentStudyCIS611Final Report.pdf"> Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study Independent Study with Nick White (Now in FaceBook and The First Prize Winner of 2016 Senior Project) Useful Machine Learning Tutorial Sites: Keras for Image Processing/Text Processing with Deep Learning: //a href="https://keras.io/guides/"> Keras Machine Learning For Your Own Advanced Study Lab Submission Instructions: 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof. 2. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. 3. If You did Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front Page of Your Report in Bigger and Bold Font ! Useful Lab Helpers: Useful Tools: //a href="https://swagger.io/"> Swagger API Tool for REST API Developments Useful Big Data Analytic Tools Choose your System/Tool/Platform to Set Up and Get Used to: Machine Learning in Python: Natural Language Processing (NLP) for Text Preprocessing/Big Data Analytics Library in Python SpaCy: //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing Lab Assignments: The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! Wait For the Lab Submission Links Are Created on Blackboard for Each Lab Lab0: Learning Python -- Due by the End of the Second Friday of the Semester Python Data Science Platforms: See More Python Platforms Above or Lab1 Set Up Guides Below to Choose for Data Science //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) Installation Guides for MySql Server: //a href="https://www.mysql.com/downloads/"> MySql Download //a href="https://dev.mysql.com/doc/refman/5.7/en/creating-database.html"> How to Create MySQL Database If You Want to Use MS SQL Server, Installation Guides for MS SQL Server: See the Announcement Section of CIS430/530 Database Systems and Processing below for Account Creation for Microsoft Azure site for Free Download of MS Visual Studio and MS SQL Sever. //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430.html#Announcement"> How to Create MS Azure Potal Site Account See the Lab Section of CIS430/530 for Installation Instruction of MS Visual Studio and MS SQL Sever. //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430.html#Lab"> How to Download and Install MS SQL Server If you have an issue with your own computer, you can borrow a Laptop from the university Tech center. Or you can set up on the Azure Cloud or Amazon Cloud. There are the general CS computer labs in the Fenn Hall at the first floor (if they are open over the pandemic). However, any computer lab won't allow you to download and set up your own platform with a database server. The subjects of Big data are advanced and new, the course requires to set up a new system like MongoDB or Hadoop in your own system. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. If you have an issue with your own computer, you can borrow a Laptop from the university Tech center. Or you can set up on the Azure Cloud or Amazon Cloud. There are the general CS computer labs in the Fenn Hall at the first floor (if they are open over the pandemic). However, any computer lab won't allow you to download and set up your own platform with a database server. The subjects of Big data are advanced and new, the course requires to set up a new system like MongoDB or Hadoop in your own system. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font ! Always Follow the Deadline of Each Lab Assigned on the Class Blackboard. The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. Please Identify Your Course When You Ask Me in Email ! Lab 1 on Web Data Processing for Information Extraction //a href="CIS593_Lab1_FAQs.pdf"> FAQs for Lab1 on Information Extraction from Webpages //a href="https://www.infoplease.com/homework-help/history/collected-state-union-addresses-us-presidents"> Infoplease site of State Union Addresses of US Presidents //a href="https://www.infoplease.com/homework-help/us-documents/state-union-address-john-adams-december-3-1799"> Correct page of Address of John Adams December 3 1799 Important Notes: 1. Part 2 (on Combining all the address texts in one text file) Is Required For CIS593 Students. Part 2 Is NOT Required for CIS492 Students but It is for extra credit for CIS492 Students. 2. If there is a Link that Does NOT Have Any Web Page Contents, Add NULL Values for the corresponding Columns for the link 3. Do not Assume that every sites has an identical URL format. This is semi-structured data. Nothing is regular in Big data. For the irregular parts, use regular expressions or xpath as neccessary. There are two ways to do Information Extraction from Webpages. Either Method is fine for Lab1 ! Method 1: Webpage as Semi-Structured HTML DOM Tree Using XPATH -- Extra Credit !! Method 2: Webpage as Unstructured Text Using Parsing API like Beautiful Soup Python Setup Guides for Labs: //a href="https://pip.pypa.io/en/stable/quickstart/"> pip Python package Installation Guide //a href="Setup Guide for Anaconda Python and Jupyter Notebook.pdf"> Setup Guide for Anaconda Python Framework and Jupyter Notebook IDE (by TA Hemal Paneliya) //a href="Setup Guide for Anaconda Python and SpiderIDE.pdf"> Setup Guide for Anaconda Python Framework and Spider IDE and Debugger (by TA Durga Dasepalli) //a href="HowtoConnectPythonJupyterToSQLServer_pyodbc.pdf"> How to Set Up Python Jupyter Notebook to Connect SQL Server using pyodbc Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Database Server Set up is needed for the Labs. You can use any SQL Server -- MySQL, MS SQL Server, or any Database Server See Lab0 Section Above for more instructions or See the step-by-step installation guides in CIS430/530 Lab Section Below //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> CIS430/530 Lab Section //a href="Setup_MSSQLSever.pdf"> How to Download and Install MS SQL Server Note that You don't Need to Do Any Extra Set Up to Use XPATH in a JavaScript in a Client side Codes for HTML DOM Processing which will be executed by your Webbrowser. All the Recent Web Browsers Have XPATH Features in their Debugger by Default (Since 2017). You Need To Set UP DOM and XPATH for Any Server-side Scripts/Languages for Applications such as Lab1. //a href="Installation guide for Python lxml in Pycharm IDE.pdf"> Installation Guide for lxml with pip Python package for DOM with XPath and pycharm IDE for debugging Lab1 Implementation Guides: Note that the Code Examples of the Lab1 Guides below Do NOT Contain the Complete Codes to Be Executed. These are ONLY for the guides for Lab1. The Environment Configuration and the Versions of APIs Varies. Do Not Copy the Entire Codes Blindly to Do Your Lab1 since it will not work depending on the version of your python and setting up for DOM and XPath. Code Examples with DOM and XPATH: //a href="CIS593_Lab1_Guide_InfoExtraction_WebPageEmery_Partial.pdf"> Example of Lab1: General Guide in Python with DOM and XPath, and pyodbc for database operations for Information Extraction //a href="CIS593_Lab1_Guide_Web_InformationExtractionDOMXpath_A.pdf"> Example of Lab1: Guide in Python with DOM and XPath in lxml for Information Extraction //a href="CIS469_569_Lab1_GuideWebScrappingDOMXpathwithPHP.pdf"> Example of Lab1: Guide in PHP with DOM and XPath for Webscapping //a href="CIS593_Lab1_WebscrapDOMXpathGuideLinux_CFord.pdf"> Example of Lab1: Guide in Python on Linux for Webscapping Code Examples with Beautiful Soup Text Processing API: //a href="CIS469_569_Lab1_GuideWebScrapping.pdf"> Example of Lab1: Guide in Python with Beautiful Soup for Information Extraction DOM with XPATH Set Up in Any Script/Programming Languages for Server Side Applications Note that You don't Need to Do Any Extra Set Up to Use XPATH in your Serverside JavaScript in NodeJS. All the labs of CIS492/593 are Server-side data processing in a server-side script/programming language with file I/O and DOM parser with Xpath. They are NOT a clientside Javascript executed by your Webbrowser. For the Labs to Write an Application Server in this course, You Need To Add DOM Parser related APIs for Your Scripts/Programming Languages To Do the following steps: to Call a DOM Parser API to Build a DOM Tree and to Use XPATH Methods for Retrival to Extarct information You need in the Application Codes as in the Examples below. Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use. //a href="http://docs.python-guide.org/en/latest/scenarios/scrape/">Setting Up lxml parser for DOM with XPath for Python //a href="NodeJSSetupXPathXMLInJS_2.pdf"> Node.JS XPath Setup Guide and XML Parser in JavaScript based Node JS //a href="https://github.com/goto100/xpath"> XPath Setup Guide for JavaScript //a href="https://www.npmjs.com/package/xpath"> XPath Setup with npm in JavaScript for Node JS Set Up DOM with XPath for HTML and XML: //a href="HowtoDebugHTMLJAVAScriptXPath.pdf"> How to Debug HTML with JavaScript, DOM, XPath //a href="https://search.yahoo.com/search?fr=mcafee&type=E211US1289G0&p=how+to+inspect+element+in+chrome"> Learn How to Inspect Element in Chrome to Debug HTML with for DOM and XPath JavaScript: //a href="https://www.w3schools.com/js/js_debugging.asp"> JavaScript Debugger //a href="https://www.w3schools.com/xml/xpath_intro.asp"> XPath Examples in Javascript //a href="https://developer.mozilla.org/en-US/docs/Introduction_to_using_XPath_in_JavaScript">MDN Site for DOM with XPath in JavaScript //a href="https://github.com/goto100/xpath"> DOM with XPath Setup Guide for JavaScript //a href="https://github.com/goto100/xpath/blob/master/docs/xpath%20methods.md"> XPath Methods //a href="NodeJSSetupXPathXMLInJS.pdf"> Node.JS XPath Setup Guide //a href="https://www.npmjs.com/package/xpath"> npm Node.JS XPath Setup Guide //a href="https://chrome.google.com/webstore/detail/xpath-helper/hgimnogjllphhhkhlmebbmlgjoejdpjl"> Google Crome XPath Helper Setup Python: //a href="http://docs.python-guide.org/en/latest/scenarios/scrape/">Setting Up DOM with XPath for Python //a href="https://www.crummy.com/software/BeautifulSoup/bs4/doc/"> BeautifulSoup Documentation for HTML, XML for Python //a href="http://web.stanford.edu/~zlotnick/TextAsData/Web_Scraping_with_Beautiful_Soup.html"> Webscapping with DOM, XPath for Python with BeautifulSoup (From the Stanford Class) //a href="https://github.com/helloitsim/InstAnalytics/blob/fdec9d27896f4b6e61acd9dc56134ad556144057/InstAnalytics.py"> Python XPath Guide Automatic Table Creation in a SQL Server After installing a SQL server (See the step-by-step installation guides in //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> CIS430/530 Lab Section Add the codes in pyODBC as below for a Connection for the Anaconda Python Framework to Connect to Your SQL Database Server For a connection with SQL Server with servername and database name using pyodbc in Python: conn = pyodbc.connect('Driver={SQL Server};Server=YOUR_SQL_SERVERNAME\SQLEXPRESS;Database=YOUR_DATABASE_NAME;Trusted_Connection=yes;') pyodbc(Open Database Connectivity) to Connect MS SQL Server in Python: //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-1-configure-development-environment-for-pyodbc-python-development?view=sql-server-ver15"> Installation pyodbc to Connect to MS SQL Server in Python //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-3-proof-of-concept-connecting-to-sql-using-pyodbc?view=sql-server-ver15"> pyodbc to Connect to MS SQL Server in Python //a href="https://datatofish.com/how-to-connect-python-to-sql-server-using-pyodbc/"> How to Connect to MS SQL Server in Python //a href="https://stackoverflow.com/questions/33725862/connecting-to-microsoft-sql-server-using-python"> pyodbc: Examples of ODBC in Python //a href="http://www.sqlines.com/oracle/datatypes/clob"> Column Data Types to store a Large Text data to Create a Table with in MS SQL Server or other database server How to Create a Table in a SQL Server from CSV/TSV Text files //a href="https://docs.microsoft.com/en-us/sql/t-sql/statements/bulk-insert-transact-sql?view=sql-server-2017"> How to Create a Table from a File with Bulk Insert with MS SQL Server //a href="https://stackoverflow.com/questions/14330314/bulk-insert-in-mysql"> How to Create a Table from a file with Bulk Insert in MySQL Server //a href="FullBulkImportAWIDSQLServer.sql"> Example of a Script to Create Multiple Tables from different files with Bulk Insert in MS SQL Server Earlier Features to Handle Big Data in Relational Database Server Advanced Data Types: BLOB (Binary Large Object) or Text/CLOB (Character Large Object) in MySQL or MS SQL Server How to Create and Insert to a Table with a Column of Large Text Data or Image Data in a Relational Database Server: //a href="https://www.mysqltutorial.org/mysql-text/"> Text Data Type in MySQL //a href="https://www.tutorialspoint.com/What-is-TEXT-data-type-in-MySQL"> What is TEXT data type in MySQL //a href="https://www.tutorialspoint.com/what-is-the-difference-between-blob-and-clob-datatypes "> Difference between blob and clob datatypes //a href="https://dev.mysql.com/doc/refman/8.0/en/blob.html "> Blob Data type in MySQL //a href="https://stackoverflow.com/questions/10729824/how-to-insert-blob-and-clob-files-in-mysql"> how to insert blob and clob from files in mysql //a href="https://stackoverflow.com/questions/752277/cannot-insert-string-into-mysql-text-column "> how to insert text column in mysql For Those Who Want to Use CLR Table Function to Create a Table in SQL Server -- This is NOT For CIS492/593 //a href="HowtoSETUPUDTASPNET.pdf"> How to Set Up ASP.NET with SQL Server //a href="HowToDebugSQCLRNOExamples.pdf"> How to Debug CLR UDF, CLR UDT, CLR TVF Lab 2 on JSON Data Processing Notes and Corrections: Note that you need to transform business.json file only for Lab2, not all of 5 json files from the Yelp site. Yelp Business Data Set for Lab2: Yelp data set and documentation for JSON File Structures You can directly download the most recent data sets from Yelp site: Yelp Full Data Set Also Avaliable in the Big data Lab below: //a href="https://eecs.csuohio.edu/~sschung/cis612/YelpDataSets.zip"> Zip file for 5 JSON files from Yelp data Challenge 2017 (Or Download directly from the Yelp site below for 2020 data sets!) Note: You Have to Use 7-Zip to unzip the zip file. Some other compression software might not be able to unzip correctly. The full JSON data files from the Yelp site might have a few incorrect JSON syntax detected in the data file or invalid line feed, which is common. Correct them before processing if detected. See FAQs for how to correct invalid JSON data Invalid JSON format handling: If there is any invalid data format is found in the input file, you can change it to the correct JSON format. For example, $$ in the “Price Range” key value pair in your input json file, the value $$ is not in quotes and this will cause to fail. Suggested Solutions: Replace $$ with 2 (meaning the price level is 2 in scale 1 - 5) in the file and try to parse the corrected file in your program. If , is missing between objects, add it. The given OneBusiness.json file and business100.json file in JSONData.zip are the corrected files. Lab 3 on Twitter Logging Data to NoSQL Database MongoDB (Extra Credit for This Semester) //a href="TwitterLoggingStructureJSON.pdf">Twitter Logging Structure in JSON Twitter API 1.1 is depreciated and wont be available for new developers: //a href="https://twittercommunity.com/t/deprecation-announcement-removing-compliance-messages-from-statuses-filter-and-retiring-statuses-sample-from-the-twitter-api-v1-1/170500"> Twitter API 1.1 is depreciated The New Structure for the 2.0 API for tweets: //a href="https://developer.twitter.com/en/docs/twitter-api/data-dictionary/object-model/tweet"> the New Structure for the 2.0 API for tweets //a href="FAQs_Lab3_3_TwitterDataCollection.pdf"> FAQs to Get Twitter Developer's Account Note That the Recent Changes of the Twitter Site Seem to Affect Their Response Time to Process Applications for Developer's Account. Due to the Delay on the Twitter Site Response Time, You Need to Apply ASAP to Collect the Twitter Streaming Data in Time Twitter Stream Data Collection Collect at least 10,000 Tweets Talking about either One of the Following Topics of Your Choice. Add More Related Keywords As Needed to Your Chosen Topic to Collect As Many Related Tweets Possible: Suggested Topics: 1. Any New Major Movie or Product That Was Released Recently if Any (For example, IPhone - IPhone 12, IPhone Mini) Or 2. Any Major News (For example, Russian Invasion to Ukraine) Or 3. President or Any Person of Interest, or Any Two Candidates in an Election Or 4. Covid, Corona Virus, Covid-19 Or 5. Any Topics of Your Interest as long as there are big enough to collect more than 20,000 Tweets Twitter Data Collection Setting Up: (The Twitter Data Collected will be used for Lab3 and Can Be Used For Your Final Project later) NOTE that to collect the Twitter stream data in real time, you need to apply for their developer’s account in the Twitter Deveoper's site and get a permission to get credentials for a token and keys. This process usually takes 3-4 days or one week for Twitter to respond. You HAVE TO start your application ASAP. Don’t wait until the last day. See the example project and sample codes for the step by step procedures for this. Read everything posted in the links below. Apply a Twitter Developer's Account ASAP for Twitter Stream Data Collection ! Start ASAP Since it Will Take a Week to Obtain a Twitter Developer's Account !! For Your Twitter Developer's Account Application, Choose the most Common Account type. Do not choose an Academic Reserach Account (You Will Be Asked More Questions). Do NOT Blindly Copy Those Sample Answers in the Class Webpage for Your Answers ! Rephrase/Modify Them in Your Words For Your Case. //a href="Example of Answers to Obtain Tweeter Developer.pdf"> See Examples of the Answers for the Questions from Twitter to Get a Twitter Developer's Account //a href="SampleAnswersForStudent_TwitterDeveloperAccountApplication.pdf"> (This is For an Academic Research Account, which You Don't Need to Apply) See Sample Answers for the New Questions from Twitter to Apply Twitter Developer's Account //a href="FAQs_Lab3_3_TwitterDataCollection.pdf"> FAQs to Get Twitter Developer's Account For Your Twitter Account Application For the Project Site if Asked, Provide Your Lab3 Specification above. you Can also Provide the CIS612 Project Site and Research Project Description for Big Data and Data Scientist //a href="https://eecs.csuohio.edu/~sschung/cis612/CIS612Project.html"> CIS 612 Project Site to Provide in Your Application for Twitter Developer's Account //a href="Faculty_Led_Project_Description__Social MeadiaSentimentAnalysisSystem_SunnieChung.pdf"> To Answer with Sample Research Project Description for Big Data and Data Scientist Twitter API 1.1 is depreciated and wont be available for new developers: //a href="https://twittercommunity.com/t/deprecation-announcement-removing-compliance-messages-from-statuses-filter-and-retiring-statuses-sample-from-the-twitter-api-v1-1/170500"> Twitter API 1.1 is depreciated The New Structure for the 2.0 API for tweets: //a href="https://developer.twitter.com/en/docs/twitter-api/data-dictionary/object-model/tweet"> the New Structure for the 2.0 API for tweets How to Collect a Twitter Stream to a JSON File then Insert to MongoDB //a href="TwitterLoggingStructureJSON.pdf">Twitter Logging Structure in JSON Note some Tweets Don't have the Retweet part of info. Try to Collect Retweets Together. How to Get Twitter Stream in JSON file: //a href="CollectingTweetStreamingDataintoJSONfileWithPython_Sp2022.pdf"> Step by Step Guide on How to Get Twitter Stream in JSON file in Python (After the Tweepy API Version 4.0 as of Spring 2022) NEW POST !! by TA Yixi Luo **************** //a href="Fetching tweets data from Twitter with Python.pdf"> Step by Step Guide on How to Get Twitter Stream in JSON file in Python (Before The Tweepy New Version 4.0) How to Get Twitter Stream in JSON file then Insert into MongoDB: //a href="StepByStepInstructions_CollectingTwitterStreamFromDeveloperAccount.pdf"> Tutorial on How to Get Twitter Stream to Insert into MongoDB in Python As of 2021 Before the New Tweepy Version 4.0 Note: This tutorial uses last year (2021)'s Tweepy Streaming Library. It has been upgraded to a new version as of early 2022. See The Tutorial for the New Version of Tweepy Streaming Lib Above Blindly copy and paste of the codes in this tutorial wouldn't work because of the new version of Tweepy Streaming Lib //a href="FetchingTweetsfromTwitteronHDFSMongoDB.pdf"> Other Related Documentations and Examples from Twitter sites and Python for Twitter If You Failed to Get a Twitter Developer's Account, Use this Twitter Raw data set in this site for This Lab //a href="https://www.kaggle.com/code/prathamsharma123/clean-raw-json-tweets-data"> Kaggle: Clean Raw JSON Tweets Data site //a href="RawJsonTwitterData.zip"> RawJsonTwitterData.zip The Recent Changes of the Twitter Site Seem to Affect Their Response Time to Process Applications for Developer's Account. Due to the Delay on the Twitter Site Response Time, Lab3 on Twitter Logging data Has Been Replaced by Lab3_1 on Yelp Data Processing. Lab3_1 with Yelp Data Set: Semi-Structured Data Processing with Semi-Structured Database Server MongoDB: 1. Creating a Semi-Structured Database - MongoDB Collections in a MongoDB Server for business.json and review.json files from the Yelp site 2. Writing Aggregation Pipelinig //a href="CIS593_Lab3_Yelp_MongoDB_BusinessDataOnly_ExtraCredit_Join.pdf"> Lab3_1 on Writing Aggregation Pipelinig on Mongo DB Collections from json data files from Yelp site Make Sure to Use All the Documents in business.json to create a Collection named business for this Lab Data Sets for Lab3_1: //a href="YelpDataSets.zip"> Zip file for 5 JSON files from Yelp data Challenge 2017 (They may not be in valid JSON format) or Download directly from the Yelp site below for 2020 data sets! Note: If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip. You can directly download the most recent data sets from Yelp site: //a href="https://www.yelp.com/dataset"> Yelp Site for Data Sets //a href="JSONData.zip"> //a href="Lab3JSON.zip"> Zip file for One business data and 100 busineess from Lab2 (in valid JSON format) If you have invalid JSON file problem, the JSON file might need to be corrected through the JSON validator to detect the errors first then correct the JSON syntax to be imported. The data files in the Zip file are all coming from the yelp site. They sometime have the incorrect JSON syntax, which need to be corrected to make it work. //a href="FAQOnLab3_YelpMongoDB.pdf">FAQs On Lab3_1 on Yelp Data with MongoDB See the example of how to correct an invalid json file below. Scroll down for the part. //a href="CIS593_Lab2_JSON_FAQs.pdf"> FAQs for JSON Processing in General from Lab2 FAQs More FAQs on Lab3 on MongoDB: Q: Can we use the 2019 dataset (JSON) from yelp website for our Lab 3_1? Or we should only use 2017 JSON data? And Do we need to create CSV for each query as well for the count result? A: Either data set is ok. Some people have a Zip error with the 2017 data set. If then use the 2019 data set from the Yelp site. Conversion to CSV file is not required for This Lab3_1. Lab3 Mongo DB Guides: Mongo DB Setup: MongoDB Installation: //a href="https://docs.mongodb.com/manual/installation/"> MongoDB Download and Installation //a href="https://docs.mongodb.com/manual/tutorial/install-mongodb-enterprise-on-windows/"> MongoDB Installation Guides //a href="https://www.mongodb.com/download-center#production"> Full Menu for MongoDB Download //a href="https://www.mongodb.com/download-center#enterprise"> MongoDB Enterprise Download //a href="https://www.mongodb.com/download-center#compass"> MongoDB GUI Client Compass Download //a href="https://docs.mongodb.com/manual/"> Full Menu of MongoDB Manuals for Getting Started //a href="https://docs.mongodb.com/mongodb-shell/#mongodb-binary-bin.mongosh"> MongoDB Cilent: Mongo Shell (mongosh) to start //a href="https://docs.mongodb.com/mongodb-shell/connect/#std-label-mdb-shell-connect"> MongoDB Shell to Connect //a href="https://docs.mongodb.com/manual/tutorial/write-scripts-for-the-mongo-shell/"> MongoDB Cilent: Writing Scripts for Mongo Shell //a href="https://docs.mongodb.com/drivers/pymongo/"> How to Install pymongo driver and Connect to MongoDB Server in Python application as a client See MongoDB Lecture Notes Section for MongoDB CRUD Queries and more details //a href="https://docs.mongodb.com/guides/server/import/"> How to Import Data file to MongoDB //a href="https://docs.mongodb.com/manual/reference/program/mongoimport/"> Mongo import //a href="MongoDBQueryExamples.pdf"> Sample Runs of Mongo DB Queries //a href="https://www.mongodb.com/compatibility/json-to-mongodb"> How to Import a json file to MongoDB //a href="https://www.mongodb.com/languages/python"> How to Use MongoDB in Python for CRUDE: DB/Collection Creation, Insert, and Querying //a href="https://www.geeksforgeeks.org/how-to-import-json-file-in-mongodb-using-python/"> How to Import a JSON file to MongoDB in Python //a href="https://stackoverflow.com/questions/70089317/how-to-do-a-word-count-in-mongodb"> Example of MongoDB Aggregation Pipeling for Word Count //a href="https://stackoverflow.com/questions/42941682/storing-json-data-into-a-variable-using-python-when-inserting-into-mongodb"> How to Save MongoDB Query Results into a variable //a href="https://stackoverflow.com/questions/20769621/saving-the-result-of-a-mongodb-query"> How to Save MongoDB Query Results Lab 4: Text Analytics (Text Mining) with Information Retrieval and Natural Language Processing Methods Lab 4 on Document Vectorization with TF-IDF for Document/Webpage Clustering for Categorization -- (For Those Who Have Already Taken CIS660, Choose Lab4_3 as Lab4 !) Note that Building Inverted Index Is NOT Required for Lab4. So Document Vectorization Based on TF-IDF Can Be Done with the Frequncy Count Only with Stopword Removal without Document Fequency Important Changes for Lab4: So Document Vectorization Based on TF-IDF Is Required with the Frequncy Count (without Document Fequency) Only with Stopword Removal 1. Building an Simplified Inverted Index in a SQL Server is an Extra Credit (50%) for Lab4 For Inverted Index to Build, You Can Simplify to One Table with (Term, Doc#, TermFreq) 2. Building a Full Inverted Index Either in a SQL Server or MongoDB is an Extra Credit (100%) as in the Lecture Note as below: Dictionary table (Term, TotalDocsFreq, TotalCollectionFreq) and Posting Table(Term, Doc#, Term_Freq) 3. Build Inverted Index with NLP Pipelining (150%) for each Sentence to Extract Context Aware Information with POS or/and NER Tagger and Store them either in SQL Server or MongoDB for Retrieval Later You Can Choose to Extend One of the Labs Below with a Big Collection of Documents as a Final Group Project ! Extra Credit Lab 4_3 on Information Retrieval Methods for Content Based Document Search Engine //a href="https://eecs.csuohio.edu/~sschung/CIS593/Text_Mining_ConstructingInvertedIndex.pdf"> Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example) Text Preprocessing Library in Python SpaCy: //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing Lab 5 on Classification with Machine Learning: Requirements for Lab5: - Repeat Classification with at least two Different MLs - For Each ML, find the best input parameters to generate the best performing model For Example, with SVM, try with different kernel functions to find the best fit(model) for the given data Optional for extra credit (10%) - K-folder Cross Validation with K = 5 or 10 For Lab 5, You Can Choose to Use Any of the Following Data Sets: Note that the Data files and Data Description Files are text files, you can open them as a text file with any text editors like Notepad or Notepad++ 1. Adult Profile Data Set to Predict Income Level < 50k or NOT from Adult Data set //a href="http://archive.ics.uci.edu/ml/datasets/Adult"> UCI Site for Census Adult Profile Data Set and Data Description for Binary Classification 2. Wine Data Set with either DT or NN to Predict Wine Quality in Scale 1 - 10 //a href="https://archive.ics.uci.edu/ml/datasets/Wine+Quality"> UCI Site for Wine Data Set and Data Description for Multi-Class Classification 3. Patient Data Set for Breast Cancer Prediction for Binary Classification to Predict whether the Patient has a Breast Cancer or not ML Algorithms -- ANN or SVM Reqire Data Preprocessing: Tutorial for Data Preprocessing Normalization and One Hot Encoding for ANN or SVM //a href="https://scikit-learn.org/stable/modules/preprocessing.html#preprocessing-categorical-features"> Categorical Data transformation Methods with Binarization (One Hot Encoding) //a href="https://visualstudiomagazine.com/articles/2013/07/01/neural-network-data-normalization-and-encoding.aspx"> Data Preprocessing Methods for ANN or SVM You Can Choose Your Own Data Set to Extend Lab5 for Final Group Project ! Data Set Repository for Classification: //a href="https://www.kaggle.com/datasets"> Kaggle Data Set Repository //a href="http://archive.ics.uci.edu/ml/datasets.php"> UCI Data Set Repository for Classification //a href="https://healthdata.gov/search/type/dataset"> Health Data Set Repository for Classification //a href="https://dev.socrata.com/"> Socrata Open Data API Extra Credt Lab 5_1 on Sentiment Analysis with Machine Learning for Classification (Not Required for Everyone): Extra Lab 5_1 on Sentiment Analysis with Machine Learning for Classification: Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set Classification Goal: 1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review. 2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star) Useful Big Data Analytic Tools Choose your System/Tool/Platform to Set Up and Get Used to: Python Analytics Tools and Tutorials : Machine Learning : Other Machine Learning Platforms: Basic R Tutorials : Extra Credit Lab 6 on MapReduce and Hadoop (HDFS): Extra Credit Lab (Not Required for Everyone) Hadoop Set Up Instruction Sites: //a href="http://hadoop.apache.org/docs/r1.0.4/single_node_setup.html"> Hadoop single node setup //a href="http://hadoop.apache.org/docs/r1.0.4/cluster_setup.html"> Hadoop Cluster setup //a href="http://hadoop.apache.org/docs/r1.0.4/mapred_tutorial.html"> Map Reduce Tutorial on Hadoop Lab Guides: Newest on the Top //a href="Lab4_1InstallHadoopDataNodeError.pdf"> Hadoop Installation: How to Fix When Data Nodes are not Running (2018) //a href="https://stackoverflow.com/questions/11889261/datanode-process-not-running-in-hadoop"> Help Site on How to Fix When Data Node are not Running (2018) //a href="HowtoExecuteMapReduceinEclipse.pdf"> Procedure to How to Execute MapReduce in Eclipse to Run a Wordcount Job (2018) //a href="Lab_4_1_Alex_Chengelis.pdf"> Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 (2017) //a href="CIS 612_Lab4_1_HadoopSetUpHalley.pdf"> Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 - Permission issue (2017) //a href="Lab4_1_CIS612_HDFSInstall_MacSonal.pdf"> Procedure to Install Hadoop and Run a Wordcount Job on Mac (2017) //a href="CIS_612_Lab_4_1HadoopWordCount_Danielle"> Guideline4 for Passwordless SSH for Lab4_1 (2016) //a href="HadoopWordCountJamesTench_2015.pdf"> Guideline2 for Lab4_1 (2015) //a href="GuideforBigDataProcessingLab4ByRyanChesla.pdf"> Guideline1 for Lab4_1 (2014) A good Instruction Site for Installing and Running Hadoop //a href="Setting-up-Hadoop-made-easy.pdf"> Installing Hadoop If you have a trouble installing Hadoop from the above site with not seeing the name node, You need to delete the temp files created in standalone mode and reformat the namenode. For setting up a passwordless SSH, see Ganesh's Lab4 below as well. //a href="HadoopAssignment4Ganesh.pdf"> Installing Hadoop and running a wordcount job by Ganesh VAVILAPALLI //a href="https://www.youtube.com/watch?v=MoKW5eY5yVY"> Video for Installing Hadoop shared by Prashant Patel //a href="TroubleShootingTipswithHadoop.pdf"> Trouble Shooting Tips for Installing Haddop on VM //a href="Lab_MRHadoop_Um.pdf"> Guideline3 for Lab4_1 on Mac (2013) For Those who Want to Set Up Your HDFS Cluster On EC2 Amazon Cloud, See the Cloud Section at the end of the Class Lecture Note Section for a Student Account. Supporting Contents for Labs : More to Come ! For Installation Guides: //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430.html#Lab"> SQL Server Installation Guide site Read the Installation Guides FIRST before starting downloading ! See the guide site for More details ! How to Create a Web Application with Java Based Application Server with MS SQL Server: Amazon Cloud: //a href="https://aws.amazon.com/rds/">Amazon RDS (Relational Database Service) //a href="http://aws.amazon.com/ec2/">Amazon Elastic Cloud Computing (EC2) for Web service //a href="http://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/TUT_WebAppWithRDS.html">How to Create Amazon Virtual Hosting (EC2) with a Web Server and Amazon Database Server (RDS) Microsoft Cloud Azure: //a href="https://azure.microsoft.com/en-us/">Trial account for Microsoft AZURE Cloud //a href="MicrosoftAzureTables.pdf">How to Create/Retrieve a Table in Microsoft AZURE Cloud //a href="https://code.msdn.microsoft.com/Fix-It-app-for-Building-cdd80df4"> Sample Project on Microsoft AZURE Cloud //a href="http://www.asp.net/aspnet/overview/developing-apps-with-windows-azure/building-real-world-cloud-apps-with-windows-azure/unstructured-blob-storage">How to Create/Retrieve BLOB data in Microsoft AZURE Cloud //a href="http://www.davidchappell.com/Azure_Services_Platform_v1.1--Chappell.pdf">AZURE Cloud //a href="windows_azure_sql_database_tutorials.pdf">How to Create a SQL Database Server in Microsoft AZURE Cloud //a href="http://www.tutorialspoint.com/microsoft_azure/index.htm">Tutorial for MS AZURE Cloud Useful Resource Sites: |
| Class | Chapter / Topic / Specific Objectives / Activities |
| 1 |
|
| 2-5 |
|
| 5-6 |
|
| 10-11 |
|
| 11-12 |
|
| 12-13 |
|
| 13-14 |
|
| 15 |
|
| 6-7 |
|
| 16 |
//a href="ENACh13final-Disks-FileStructure-Hashing.pdf"> Project Presentation |
==> Completion of Homeworks/Labs is required for obtaining a passing grade.
| This
is a tentative scale and |
Letter |
Quality Points |
|
||
| A |
> 93% |
A: Outstanding (student's performance is genuinely excellent) | |||
| A- |
90% - 93% |
||||
| B+ |
87% - 90% |
||||
| B |
82% - 87% |
B: Very Good (student's performance is clearly commendable but not necessarily outstanding) | |||
|
|
B- |
80% - 82% |
|||
|
|
C |
75% - 80% |
C: Good (student's performance meets every course requirement and is acceptable; not distinguished) | ||
| D | 65%-75% | D: Below Average (student's performance fails to meet course objectives and standards) | |||
|
|
F |
<65% |
F: Failure (student's performance is unacceptable) | ||
|
ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106. |
Programming standards
|
|
|