CIS 492/593 and DSA 469 Big Data with Big Data Processing Systems (3-0-3) |
||
| Course Content |
|
|
|
The Final Exam on Tues Dec 9, 2025 at 4:00PM - 6:00PM In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final //a href="CIS593S24_FinalExamLinkOnly.html"> See the Links for Final Exam Lecture Note Only (Scroll Down to the Later Half of Lecture Note Sections) Topics to Focus on for the Final of CIS493/593 and DSA469 Design of AI Application System in Phases - How to Build an Intelligent System (AI Applications) with Data Processing Pipelining for Big data processing Unstructured Text Processing Techniques, NLP Data Preprocessing Methods Inverted Index TF-IDF, Cosine Similarity Context Aware Information Extraction (IE) methods of NLP Methides- POS Tagging, NER tagging, Lemmatization Data Preprocessing Methods for Classification with ML Algorithms - Decision Tree, ANN Classification with Machine Learning Algorithms: Decision Tree and Neural Network Backpropagation Algorithm NOT FOR THIS SEMESTER: Map Reduce Process on Hadoop Distributed File System One Page Hand Written Note (Each Side) is Allowed to the Final ! Printed Copy and Pasted of the Lecture Notes Are NOT Permitted as One Page Note Information on Midterm: Midterm will Be (Tentatively) on October 16 at 4:00PM - 5:45PM in Class. (It Will be Announced Ahead in Class) ! //a href="CIS593Sp23_MidtermOnly.html"> See Lecture Links Only for MidTerm (Scroll down to the Lecture Note Section to See Active links - All the Rest of Links are Disabled) The Topics to Focus On for the Midterm: Basically whichever Subjects on the Lecture Notes Covered in class! IMPORTANT WARNING !! If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !! Characteristics of Big Data and their Contents, Main Differences between Big data and Traditional Data, Characteristics of Semi-Structured Data Model, Main Differences between Semi-Structured Data and Structured Relational Data, Required Data Processing Steps in Common Big Data Applications with Semi-structure Data in Object Exchange Model(OEM) Universal Data Exchange Formats in Object Exchange Model(OEM): Three common big data formats as semi structured data: - HTML, XML, JSON - Data Model, Syntax of Each Encoding Format, All the related processing techniques -- DOM, XPath Conversion between Relation(CSV), XML, and JSON Comparison Between Semistructured Database Server and Relational Database Server for Database Management Strategies. Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining Object Relation Mapping (ORM) Strategies, Conversion between Relation(CSV), XML, and JSON NOTE that ONE Page Note is NOT allowed to the Midterm !! Exam Formats: 6-7 Main Questions with 2-3 subquestions for a Small/Short Answer Types for Probelm Solving 11. August 25, 2025: It is Required to Attend Every Class ! Your Class Attedance Will be Checked in Each Class. There Will be Random Quizzes to Check the Attendance ! Only 5 - 10 Mins Will Be Allowed for Each Quiz. Those Who Come in Class Late After 5 Mins, They Will NOT Be Given a Quiz ! 14. August 25, 2025: The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! Wait For the Lab Submission Links To Be Created on Blackboard with the Deadline to Submit Each Lab 10. August 25, 2025: Only the registered students can access the course blackboard. If you have a problem with your blackboard access, please contact the registar and CSU e-learning Tech Support to resolve the issue ! Faculty do not control your registration in the CSU Campus Systems and the course blackboard access. //a href="https://www.csuohio.edu/center-for-elearning/technical-support"> e-learning CSU Tech Support 9. August 25, 2025: TA Information: TA: Sowmya Chinthalapudi Primary Emails: 2898471@vikes.csuohio.edu Office Hours: Tues 12:00 PM – 2:00 PM Thursday 12:00PM PM – 2:00 PM Location: Big Data Lab: FH 305 or a ZOOM Meeting Zoom Info: Meeting ID: 588 581 6192 Passcode: (Send email to TA ahead to set up a time slot for a zoom or to let TA know that you are coming) If you have questions in Labs or grading your Labs, Send an email to TA to See During the TA's Office Hours or Schedule a Zoom meeting Dr. Chung's Office Hours: Tues and Thurs 1:30PM - 3:30PM Send Me Email Ahead to Set Up an In-Person or Zoom Meeting. Email: s.chung@csuohio.edu Zoom Meeting Info: Meeting ID: 859 3867 6332 //a href="https://csuohio.zoom.us/s/85938676332"> ZOOM Meeting Link 6. August 25, 2025: Lab Submission: The Output of each lab is your report in Doc file that shows your screen captures of the Executions with Each Output (each intermediate output as well as final outputs) Generated. your Report in Doc file also should expain all the platform set up, the execution steps, and copy of each source code files ) Each of your screen capture must show your results returned in YOUR SYSTEM to prove that your lab is done correctly by YOU !! Submit your Zip file that includes the followings on Blackboard for a timestamp and as a proof 1) Your Lab Report in .doc file that explains all the platform set up procedures, the execution steps, each intermediate output, final outputs, and a copy of each source code files, 2) All your Source files, and output files 3. August 25, 2025: Each Class Attendance is Required for CIS492/593 and DSA469 ! There Will Be Random Quizzes and Sign Up Sheet to Check Each Attendance of You ! 2. August 25, 2025: You can use any SQL Server for Your Database such as MySql or MS SQL Server. See Set up Guide for MS SQL Server or MySql in the Lab Section. 1. August 25, 2025: The class webpage for this semester was announced on blackboard: //a href="https://eecs.csuohio.edu/~sschung/CIS593/CIS593F25.html"> Class Webpage of DSA469/CIS492/593 Big Data Or You can Reach from Teaching Section of My Website for Big data Research Lab at: //a href="https://eecs.csuohio.edu/~sschung/"> Big Data Research Lab If you have a trouble to display the webpage correctly with MS Internet Explorer, open it with Google Chrome Check the Last Day to Add and Drop here ! //a href="https://www.csuohio.edu/registrar/academic-calendar"> University's Official Academic Calendar for the Semester and the Final Exam schedules Final Exam Schedule: Tues Dec 9, 2025 4:00p-6:00p Information on Midterm and Final Exams: Midterm will Be on Tentatively Thursday Oct 16 at 4:00PM - 5:45PM in Class. It Will be Announced Ahead in Class ! //a href="CIS593Sp23_MidtermOnly.html"> See Lecture Links Only for MidTerm (Scroll down to the Lecture Note Section to See Active links - All the Rest of Links are Disabled) The Topics to Focus On for the Midterm: Basically whichever Subjects on the Lecture Notes Covered in class! IMPORTANT WARNING !! If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !! Characteristics of Big Data and their Contents, Main Differences between Big data and Traditional Data, Characteristics of Semi-Structured Data Model, Main Differences between Semi-Structured Data and Structured Relational Data, Required Data Processing Steps in Common Big Data Applications with Semi-structure Data in Object Exchange Model(OEM) Universal Data Exchange Formats in Object Exchange Model(OEM): Three common big data formats as semi structured data: - HTML, XML, JSON - Data Model, Syntax of Each Encoding Format, All the related processing techniques -- DOM, XPath Conversion between Relation(CSV), XML, and JSON Comparison Between Semistructured Database Server and Relational Database Server for Database Management Strategies. Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining Object Relation Mapping (ORM) Strategies, Conversion between Relation(CSV), XML, and JSON NOTE that ONE Page Note is NOT allowed to the Midterm !! The Final Exam Information: In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final //a href="CIS593S24_FinalExamLinkOnly.html"> See the Links for Final Exam Lecture Note Only (Scroll Down to the Later Half of Lecture Note Sections) Topics to Focus on for the Final of CIS493/593 and DSA469 Design of Intelligent System with data pipelining of Big data processing Unstructured Text Processing Techniques, Data Preprocessing Methods Inverted Index TF-IDF, Cosine Similarity Context Aware Information Extraction (IE) methods of NLP Methids- POS Tagging, NER tagging, Lemmarization Data Preprocessing Methods for Classification with ML Algorithms - Decision Tree, ANN Classification with Machine Learning Algorithms: Decision Tree and Neural Network BackPropagation Algorithm NOT FOR THIS SEMESTER: Map Reduce Process on Hadoop Distributed File System One Page Hand Wrtten Note (Each Side) is Allowed to the Final ! Printed Copy and Pasted of the Lecture Notes Are NOT Permitted as One Page Note |
CIS492/593 Big Data for the CS Majored Students as an Advanced Elective Course and Also Offered as DSA469
DSA 469 Big Data Processing Systems is a Core Required Course of the Data Science Degree (BSDS)
|
Projects on Big Data Processing, Building an Intelligent Application (AI) with Big Data Analytics The Scope of Group Project Requirements Has Been Adjusted to: (This Adjustment May not Apply to this semester. Will Be Announced) 1. You Can Do One Person Group Project in a Small Scale. 2. a Project with a Small Scale with Some Extension of One of Lab4 - Lab4_3 3. For Those Who Have Aleady Taken CIS660, it is Required to Complete a Full Project with Bid data set and Presentation 4. Implementing Your Final Project with Clientside and Serverside of a Web Application with a Web Based User Interface Is NOT Required. It will Be Counted as Extra Credit. Final Group Project Specification and Instructions: Project Submission Instructions: Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file either on the google drive or One Drive and send the link. (Send me email for access permission for this link !) Submit a Zip file that includes: All of your presentation slides (both in .ppt and .pdf) and Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files on Blackboard by the end of Friday of your presentation week. Include the Problems/Error Encountered and Your Resolutions in Your Report One Submission Per Group Required. Submit a Zip File that Includes All the required Source Files, Input, Output Files, and Final Report (in Doc file) and Presentation Slides (in pptx). Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. Important Notes for Final Project: - 1 - 2 Person Group Project Are Allowed for the Small Class Size (~25 students); 1 - 3 Person Group Project Are Allowed for the Big Class Size (> 35 students); 2 - 4 Person Group Project Are Allowed for the Very Big Class Size (> 40 students) - One Person Final Project is Allowed. You Can Work Alone. - If Three Person Group is Allowed, Make Sure to Make the Project in a Bigger Scale - You can Change Your Project Plan/Proposal or the Details even after Your Proposal is submitted until the Deadline of the Status Report. - If You Need to Find Group Members, Use Blackboard Email to Send to the Class, Some will respond to you if they are looking for a group member. The IMDB Data links in the Project List are gone. Check here for Movie Review Data Set. //a href="https://www.kaggle.com/rounakbanik/the-movies-dataset#movies_metadata.csv"> Movie Review Data //a href="https://developers.google.com/youtube/reporting/"> Youtube Analytic Site For Those Who Have Already Taken CIS660, The Final Group Project Should Include Fully Analytic Processing Suggested Projects: For Text Analytics like Sentiment Analysis or Opinion Analysis: NLP Techniques - POS, NER Tagging, Bi-Gram Handling Are Required for Preprocessing. For Document Categorization: by Constructing TF-IDF Vectorization. Inverted Index Building Will Be Plus but Optional. Building Word2Vec Embeddings for a Collection of Documents/Webpages with Training Set Generation in Skip Gram Model For Other Types of Projects, the Proposal is Required to Be Approved to Meet the Complexity of Final Project Extra Credit Final Project: Final Group Project Time Line: (Tentative) The Exact Deadline of Each Submission Will Be Posted on teh Class BlackBoard Group Project Proposal Due by November 7 ! Group Project Status Report Due by November 16 ! Group Project Presentation Either on Dec 2 or 4 ! Final Group Project Report Due By Friday Dec 5 ! If the Size of the Class is Too Big, 3-4 Person Project Group Will Be Allowed As Long As the Project Scale is Big Enough for 3-4 Persons IMPORTANT Submissions for Group Project Task 1: Group Project Proposal (Plan) Submit minimum 3 Page Group Proposal on Blackboard with List of Group Memebers Group Project Proposal (Plan) Should Include Brief Descriptions on: 1. Description of Big Data in Size and Format, Data Collection Plan, 2. Goal of Your Big Data Analytic Project with What Kind of Intelligent Analytic Funtionality, Features of Your AI/Big Data Analytic Application 3. Big Data Processing Plan, Methods 4. Investigate on Platform/Systems/Tools/APIs to Use Task 2: Group Project Status Report Your Project Status Report Should Show the Following Tasks Done: 1. Platform Setting/System Configuration Procedure (if it is new) 2.Your KnowledgeBase Structure/ Database Design 3. Design of Big Data Processing Pipeline, Data Transformation Methods 4. Your Server Data/Database Contents in Progress, Any Intermediate Outputs in Progress Task 3: Project Presentation Project Presentation Starts from the Last Week of the Class of the Semester Project Presentation Schedule Will Be Sent To Your CSU Email for Sign Up One week Before the Presentation Read the Instructions of Project Presentation Here ! Group Project Presentation Should Include: 1. Data Description, Data Size, Data Collection Method 2. Goal of Your Intelligent Big Data Analytic Application (AI) 3. Platform Setting/System Configuration Procedures 4. System Design (Architecture) of Your AI Application in Detail 5. Raw Big Data Preprocessing Methods and Intermediate Results 6. Design of Big Data Processing Pipeline, Data Transformation Methods 7. Description of Your KnowledgeBase Structure/Database Design for Your AI Application. And Show the Contents 8. Your ML Algorithms Used, Ranking Algorithm Used for Labelling if any 9. Your ML Models and Evaluation (Accuracy) Results and Visualization of the Result, Data Marix (Structure) for Analysis/Evaluation and Visulation of the Analysis Results (for example, correlation matrix or similarity matrix) if Any 10. The Problems/Errors Encountered and Your Resolutions 11. System Demo Submit Final Project Report and Presentation By the End of Friday of The Last Class Week After Your Presentation! One Submission Per Group Required. Submit a Zip file that includes: 1. All of your presentation slides (both in .pptx) and 2. Your Group Final Project Report (in doc) The Final Project Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. The Report Should Include Platform/System Set up the Set Up Procedure /Configuration Detail of Your Platform/System/Packages, Executions Steps, all the Source Codes, Scripts, all the intermediate outputs, and final output files. Include the Problems/Error Encountered and Your Resolutions in Your Report If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. Group Project Presentation Should Include: 1. Data Description, Data Size, Data Collection Method 2. Platform Setting/System Configuration Procedures 3. System Design (Architecture) of Your AI Application or Data Analytic Goal in Detail 4. Raw Big Data Preprocessing Methods and Intermediate Results 5. Design of Big Data Processing Pipeline, Data Transformation Methods 6. Description of Your KnowledgeBase Structure/Database Design. And Show the Contents 7. Ranking Algorithm, Data Matrix (Structures) if any for Evaluation 8. The Problems/Errors Encountered and Your Resolutions 9. System Demo or Evaluation Results and Visualization of the Result < Requirements of Combining two Group projects of two related courses: Combining two Group projects of two related courses (CIS492 with CIS408 or CIS593 Deep Learning) are ok as long as the combined project focuses on both sides of the subjects for example, with Web application aspect for CIS408 and Server-side big data processing with a database server and any analytic methods with big data covered in CIS492. The data of the Deep Learning class CIS593 is images, which is a lot different from the Big data covered in this Big data class such as a large volumn of semistructured collection or text documents. Maybe training a large text documents using deep learning would be a good combined project of CIS492/CIS593 and the Deep Learning Class. For Extra Credit Projects or Contract Course Requirement for Honor Students It includes a project to build a real-life web based AI applications like: - Web Search Engine for a certain domain, for example, .csuohio.edu or .cnn - WebMD - Question Answering system like Google or Alexa but in a Webbased Interface to get a user question Or the requirements in the CIS492/DSA469 Honor Contract Course Description //a href="HonorsContractCourseDescription_CIS492_DSA469.pdf"> CIS492/DSA469 Honor Contract Course Description 1) Integrate Big Data processing system with a Web Application through building a big data application with web technologies. 2) Big Data Analytic Method with Machine Learning 3) Any Parallel Data processing technologies such as any Apache Hadoop with Google’s MapReduce based parallel data processing systems – Hive or Spark, or PySpark that is integrated with a web application through building a big data application with web technologies Project Examples of Vectorization of each document in a Training set for a Machine Learning Classifier: //a href="https://eecs.csuohio.edu/~sschung/CIS660/ProjectExample_SentimentAnalysis_YelpReview.pdf"> Project Example I: Sentiment Analysis of Yelp Business Review with Machine Learning //a href="https://eecs.csuohio.edu/~sschung/CIS660/ProjectExample_SentimentAnalysis_YelpReview_Feng.pdf"> Project Example II: Sentiment Analysis of Yelp Business Review with Machine Learning Best Senior Design Projects on Big Data Processing and Text Anaytics (Created from the Subject of CIS492/593, CIS430/530 and CIS408): 2022 - 2024: Big Data and AI Projects: 2020 - 2021: Big Data and AI Projects: 2019 - 2017: Big Data and Data Science Projects: The Best Senior Design Projects Created from CIS408 and CIS430: //a href="Cerebro Poster v3NickWhite.pdf"> The First Prize Winner of 2016 Senior Project From CIS430 and CIS408 by Nick White (Now in FaceBook), et al Best Group Projects: Selected Best Projects Will be Posted Here ! You Can Choose to Extend One of The Extra Credit Labs Below as a Final Group Project either with a Different data set or the same data set Extra Credit Lab 4_2 on Webpage Categorization by Topics Extra Credit Lab 4_3 on Information Retrieval Methods for Content Based Document Search Engine //a href="https://eecs.csuohio.edu/~sschung/CIS593/Text_Mining_ConstructingInvertedIndex.pdf"> Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example) Extra Credt Lab 5 on Sentiment Analysis with Machine Learning for Classification: Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set Classification Goal: 1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review. 2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star) More To Come Here ! Group Project Data Sources: You can choose to work on these data sets for your group project Final Project Submission Instructions: Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file on your google drive or One Drive and Send email to me and TA to share ! Submit a Zip file on Blackboard by the end of Friday of your presentation week. One Submission Per Group Required. Your Project Zip File Should includes: 1) All of your presentation slides (both in .ppt and .pdf) and 2) Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files 3)Include the Problems/Error Encountered and Your Resolutions in Your Report Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. IMPORTANT NOTE !!! If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. |
|
Lab Assignment 0: - Learning Python -- Due by the End of the Second Weekend of the Semester Basic Python Tutorials : - Set Up Your Python Platform -- by the End of the First Weekend ! Set Up For the upcoming Lab1 Set Up Guides for Important Platform with Python, XPath, Beautiful Soup, and MySQL:****************************** You can use any SQL Server -- MySQL, MS SQL Server, or any Database Server You Need To Set UP DOM and XPATH for Any Server-side Scripts/Languages for Applications such as Lab1. //a href="CIS492_593_DSA469_Lab1_Environment_Setup.pdf"> Platforms Set up Guides with Python, MySql, XPath, PyCharm Debugging IDE for Big Data Labs (by TA Dennis Risch) ******************************* See More Guides in the Lab1 Section (Far Below) Set Up Your Python Data Science Platform for Other Labs and Final Projects There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library ! See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !! Additional Installation Guide: //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) • Anaconda //a href="https://www.anaconda.com/open-source"> Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms //a href="https://docs.anaconda.com/anaconda/navigator/tutorials/"> Anaconda Tutorials • Python Anaconda Tutorial Sites //a href="https://data-flair.training/blogs/python-anaconda-tutorial/"> Anaconda Tutorial Site //a href="https://www.edureka.co/blog/python-anaconda-tutorial/"> Anaconda Tutorial Site • PyTorch //a href="https://pytorch.org/"> PyTorch Site (It can be integrated from Anaconda as well) Google Colab: //a href="https://www.dataquest.io/blog/jupyter-notebook-tutorial/"> Basic Guide for Python tools for Data Science: jupyter-notebook //a href="https://www.dataquest.io/blog/advanced-jupyter-notebooks-tutorial/"> More Basic Guide for Python tools for Data Science: jupyter-notebook • Python Scikit Learn for Common Data Science Tasks Text Preprocessing (Natural Language Processing) Library in Python SpaCy: //a href="https://spacy.io/api"> Liquistic Modules in Python SpaCy //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing • Python sklearn.cluster //a href="https://scikit-learn.org/stable/modules/clustering.html"> Python Sklearn Clustering Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Basic R Tutorials : //a href="IndependentStudyCIS611Final Report.pdf"> Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study Independent Study with Nick White (Now in FaceBook and The First Prize Winner of 2016 Senior Project) Useful Machine Learning Tutorial Sites: Keras for Image Processing/Text Processing with Deep Learning: //a href="https://keras.io/guides/"> Keras Machine Learning For Your Own Advanced Study Lab Submission Instructions: 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof. 2. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. 3. If You did Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front Page of Your Report in Bigger and Bold Font ! Useful Lab Helpers: Useful Tools: //a href="https://swagger.io/"> Swagger API Tool for REST API Developments Useful Big Data Analytic Tools Choose your System/Tool/Platform to Set Up and Get Used to: Machine Learning in Python: Natural Language Processing (NLP) for Text Preprocessing/Big Data Analytics Library in Python SpaCy: //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing Lab Assignments: The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! Wait For the Lab Submission Links Are Created on Blackboard for Each Lab Lab Assignment 0: - Learning Python and Setting Up Your Python Platform for Lab1 -- Due by the End of the Second Friday of the Semester Python Data Science Platforms: See More Python Platforms Above or Lab1 Set Up Guides Below to Choose for Data Science //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) - Set Up Your Python Data Science Platform for Labs and Final Projects by the End of the First Weekend ! There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library ! See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !! For the upcoming Lab1 Set Up Guides for Important Platform with Python, XPath, Beautiful Soup, and MySQL:****************************** You can use any SQL Server -- MySQL, MS SQL Server, or any Database Server You Need To Set UP DOM and XPATH for Any Server-side Scripts/Languages for Applications such as Lab1. //a href="CIS492_593_DSA469_Lab1_Environment_Setup.pdf"> Platforms Set up Guides with Python, MySql, XPath, PyCharm Debugging IDE for Big Data Labs (by TA Dennis Risch) ******************************* More on Setup Guides: //a href="https://pip.pypa.io/en/stable/quickstart/"> pip Python package Installation Guide //a href="Setup Guide for Anaconda Python and Jupyter Notebook.pdf"> Setup Guide for Anaconda Python Framework and Jupyter Notebook IDE (by TA Hemal Paneliya) //a href="Setup Guide for Anaconda Python and SpiderIDE.pdf"> Setup Guide for Anaconda Python Framework and Spider IDE and Debugger (by TA Durga Dasepalli) pyodbc Set Up Guide With MS SQL Server: //a href="HowtoConnectPythonJupyterToSQLServer_pyodbc.pdf"> How to Set Up Python Jupyter Notebook to Connect to a Microsoft SQL Server using pyodbc**************** Platform Set Up Guide on Mac: //a href="Installation guide for Python lxml in Pycharm IDE.pdf"> Installation Guide on Mac for lxml with pip Python package for DOM with XPath and pycharm IDE for debugging on Mac More on Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Installation Guides for MySql Server: //a href="https://www.mysql.com/downloads/"> MySql Download //a href="https://dev.mysql.com/doc/refman/5.7/en/installing.html"> MySql Download and Installation //a href="https://dev.mysql.com/doc/refman/8.2/en/tutorial.html"> MySql Tutorial //a href="https://dev.mysql.com/doc/refman/5.7/en/creating-database.html"> How to Create MySQL Database If You Want to Use MS SQL Server, Installation Guides for MS SQL Server: See the Announcement Section of CIS430/530 Database Systems and Processing below for Account Creation for Microsoft Azure site for Free Download of MS Visual Studio and MS SQL Sever. //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Announcement"> How to Create MS Azure Potal Site Account See the Lab Section of CIS430/530 for Installation Instruction of MS Visual Studio and MS SQL Sever. //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> How to Download and Install MS SQL Server If you have an issue with your own computer, you can borrow a Laptop from the university Tech center. Or you can set up on the Azure Cloud or Amazon Cloud. There are the general CS computer labs in the Fenn Hall at the first floor (if they are open over the pandemic). However, any computer lab won't allow you to download and set up your own platform with a database server. The subjects of Big data are advanced and new, the course requires to set up a new system like MongoDB or Hadoop in your own system. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. If you have an issue with your own computer, you can borrow a Laptop from the university Tech center. Or you can set up on the Azure Cloud or Amazon Cloud. There are the general CS computer labs in the Fenn Hall at the first floor (if they are open over the pandemic). However, any computer lab won't allow you to download and set up your own platform with a database server. The subjects of Big data are advanced and new, the course requires to set up a new system like MongoDB or Hadoop in your own system. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font ! Always Follow the Deadline of Each Lab Assigned on the Class Blackboard. The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. Please Identify Your Course When You Ask Me in Email ! XML/XHTML/JSON Syntax Validators: PreLab1 on HTML: Creating Your Personel Webpage Step 1: Create Your Personal Webpage to Display Information about Youself with: The Minimum requirements to Include: 1) Your Name in a big heading 2) Your Picture with any Images 3) 2 Links to Any Related Sites to You (for example, the CSU site, the CS Dept Site) 4) Two Fun Things You Like to Do (in Bullet Point for each) 5) Button with "Click Me" and onclick attribute binded to a javascript function to CHANG the Text of the Second Fun Thing You Like to Do to SOMETHING DIFFERENT in red color Step 2: Launch Your Personal Website on the CS Grail Linux Server where an Apathe Webserver is Running Follow the Instruction under Project 0 on My CIS408 Lab Section: For Step2, You Need to Have a Linux Account on the CS Dept Grail Linux Server ! Follow the instructions to create a Linux account in the CIS408 Project0 Section above. If you have a issue on your Linux account or the Linux Server, Send email to Our Dept System Specialest Joe Ruess at: j.ruess@csuohio.edu Lab 1 on Web Data Processing for Information Extraction //a href="CIS593_Lab1_FAQs.pdf"> FAQs for Lab1 on Information Extraction from Webpages //a href="DSA469CIS492_593_ExampleLabReportOutput.pdf"> Lab Report Example for Lab1 //a href="https://www.infoplease.com/homework-help/history/collected-state-union-addresses-us-presidents"> Infoplease site of State Union Addresses of US Presidents (As of 2025, This site blocks any http request from unknown clients (you)) //a href="https://www.infoplease.com/homework-help/us-documents/state-union-address-john-adams-december-3-1799"> Correct page of Address of John Adams December 3 1799 Another site for State Union Addresses of US Presidents that Does NOT block http request (Found by Nicholas Rinaldi) Important Notes: 1. Part 2 (on Combining all the address texts in one text file) Is Required For CIS593 Students. Part 2 Is NOT Required for CIS492 Students but It is for extra credit for CIS492 Students. 2. If there is a Link that Does NOT Have Any Web Page Contents, Add NULL Values for the corresponding Columns for the link 3. Do not Assume that every sites has an identical URL format. This is semi-structured data. Nothing is regular in Big data. For the irregular parts, use regular expressions or xpath as neccessary. There are two ways to do Information Extraction from Webpages. Method 1: Webpage as Semi-Structured HTML DOM Tree Using XPATH -- BETTER WAY (Required) to Learn For Lab1 !! Method 2: Webpage as Unstructured Text Using Parsing API like Beautiful Soup The Lab1 is to learn HTML DOM Processing using XPATH ! Set Up Guides for Important Platform with Python, XPath, Beautiful Soup, and MySQL:****************************** You Need To Set UP DOM and XPATH for Any Server-side Scripts/Languages for Applications such as Lab1. //a href="CIS492_593_DSA469_Lab1_Environment_Setup.pdf"> Platforms Set up Guides with Python, MySql, XPath, PyCharm Debugging IDE for Big Data Labs (by TA Dennis Risch) ******************************* More on Setup Guides: //a href="https://pip.pypa.io/en/stable/quickstart/"> pip Python package Installation Guide //a href="Setup Guide for Anaconda Python and Jupyter Notebook.pdf"> Setup Guide for Anaconda Python Framework and Jupyter Notebook IDE (by TA Hemal Paneliya) //a href="Setup Guide for Anaconda Python and SpiderIDE.pdf"> Setup Guide for Anaconda Python Framework and Spider IDE and Debugger (by TA Durga Dasepalli) pyodbc Set Up Guide With MS SQL Server: //a href="HowtoConnectPythonJupyterToSQLServer_pyodbc.pdf"> How to Set Up Python Jupyter Notebook to Connect to a Microsoft SQL Server using pyodbc**************** Platform Set Up Guide on Mac: //a href="Installation guide for Python lxml in Pycharm IDE.pdf"> Installation Guide on Mac for lxml with pip Python package for DOM with XPath and pycharm IDE for debugging on Mac More on Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Database Server Set up is needed for the Labs. You can use any SQL Server -- MySQL, MS SQL Server, or any Database Server See Lab0 Section Above for more instructions or See the step-by-step installation guides in CIS430/530 Lab Section Below //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> CIS430/530 Lab Section //a href="Setup_MSSQLSever.pdf"> How to Download and Install MS SQL Server Note that the Set Up Guides Above are for Server Side Applications such as Lab1 in Python or Java. You don't Need to Do Any Extra Set Up to Use XPATH in a JavaScript in a Client side Codes for HTML DOM Processing which will be executed by your Webbrowser. All the Recent Web Browsers Have XPATH Features in their Debugger by Default (Since 2017). IMPORATNT NOTE !!!! The tutorials and examples in the class Lab Section are to guide and teach the methods and techniques. Please note that they are NOT to provide you with precise coding solutions that would run without checking/debugging the changes in codes; Especially since these areas are fast changing and there are so many variations in the platforms/languages. YOU ARE RESPONSIBLE for DEBUGGING YOUR CODES !! Lab1 Implementation Guides:******************************* //a href="Lab1Guides_Links for XPath_Pyodbc_BLOBDataType.pdf"> Lab1 Guides: What to Need to Know in the Lab1 Section to Do Lab1 ************************* Note that the Code Examples of the Lab1 Guides below Do NOT Contain the Complete Codes to Be Executed. These are ONLY for the guides for Lab1. The Environment Configuration and the Versions of APIs Varies. Do Not Copy the Entire Codes Blindly to Do Your Lab1 since it will not work depending on the version of your python and setting up for DOM and XPath. How to Write Xpath in Demo in Webbrowser: (Warning: As of 2025, There are minor syntax changes in these notes: the double quotation mark is not valid in xpath operator $x() //a href="Lab1_Xpath_Syntax_WebpageDemo.pdf"> Example of Xpath Execution in Webbrowser ************************* //a href="Lab1_Xpath_WebpageDemo2.pdf"> Code Examples with DOM and XPATH: //a href="CIS593_Lab1_Guide_InfoExtraction_WebPageEmery_Partial.pdf"> Example of Lab1: General Guide in Python with DOM and XPath, and pyodbc for database operations for Information Extraction ************************* //a href="CIS593_Lab1_Guide_Web_InformationExtractionDOMXpath_A.pdf"> Example of Lab1: Guide in Python with DOM and XPath in lxml for Information Extraction ************************* //a href="CIS469_569_Lab1_GuideWebScrappingDOMXpathwithPHP.pdf"> Example of Lab1: Guide in PHP with DOM and XPath for Webscapping //a href="CIS593_Lab1_WebscrapDOMXpathGuideLinux_CFord.pdf"> Example of Lab1: Guide in Python on Linux for Webscapping Code Examples with Beautiful Soup Text Processing API: //a href="CIS469_569_Lab1_GuideWebScrapping.pdf"> Example of Lab1: Guide in Python with Beautiful Soup for Information Extraction DOM with XPATH Set Up in Any Script/Programming Languages for Server Side Applications Note that You don't Need to Do Any Extra Set Up to Use XPATH in your Serverside JavaScript in NodeJS. All the labs of CIS492/593 are Server-side data processing in a server-side script/programming language with file I/O and DOM parser with Xpath. They are NOT a clientside Javascript executed by your Webbrowser. For the Labs to Write an Application Server in this course, You Need To Add DOM Parser related APIs for Your Scripts/Programming Languages To Do the following steps: to Call a DOM Parser API to Build a DOM Tree and to Use XPATH Methods for Retrival to Extract information You need in the Application Codes as in the Examples below. Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use. //a href="http://docs.python-guide.org/en/latest/scenarios/scrape/">Setting Up lxml parser for DOM with XPath for Python ***************************** //a href="NodeJSSetupXPathXMLInJS_2.pdf"> Node.JS XPath Setup Guide and XML Parser in JavaScript based Node JS //a href="https://github.com/goto100/xpath"> XPath Setup Guide for JavaScript //a href="https://www.npmjs.com/package/xpath"> XPath Setup with npm in JavaScript for Node JS Set Up DOM with XPath for HTML and XML: //a href="HowtoDebugHTMLJAVAScriptXPath.pdf"> How to Debug HTML with JavaScript, DOM, XPath //a href="https://search.yahoo.com/search?fr=mcafee&type=E211US1289G0&p=how+to+inspect+element+in+chrome"> Learn How to Inspect Element in Chrome to Debug HTML with for DOM and XPath JavaScript: //a href="https://www.w3schools.com/js/js_debugging.asp"> JavaScript Debugger //a href="https://www.w3schools.com/xml/xpath_intro.asp"> XPath Examples in Javascript //a href="https://developer.mozilla.org/en-US/docs/Introduction_to_using_XPath_in_JavaScript">MDN Site for DOM with XPath in JavaScript //a href="https://github.com/goto100/xpath"> DOM with XPath Setup Guide for JavaScript //a href="https://github.com/goto100/xpath/blob/master/docs/xpath%20methods.md"> XPath Methods //a href="NodeJSSetupXPathXMLInJS.pdf"> Node.JS XPath Setup Guide //a href="https://www.npmjs.com/package/xpath"> npm Node.JS XPath Setup Guide //a href="https://chrome.google.com/webstore/detail/xpath-helper/hgimnogjllphhhkhlmebbmlgjoejdpjl"> Google Crome XPath Helper Setup Python: //a href="http://docs.python-guide.org/en/latest/scenarios/scrape/">Setting Up DOM with XPath for Python //a href="https://www.crummy.com/software/BeautifulSoup/bs4/doc/"> BeautifulSoup Documentation for HTML, XML for Python //a href="http://web.stanford.edu/~zlotnick/TextAsData/Web_Scraping_with_Beautiful_Soup.html"> Webscapping with DOM, XPath for Python with BeautifulSoup (From the Stanford Class) //a href="https://github.com/helloitsim/InstAnalytics/blob/fdec9d27896f4b6e61acd9dc56134ad556144057/InstAnalytics.py"> Python XPath Guide Automatic Table Creation in a SQL Server After installing a SQL server (See the step-by-step installation guides in //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> CIS430/530 Lab Section Add the codes in pyODBC as below for a Connection for the Anaconda Python Framework to Connect to Your SQL Database Server For a connection with SQL Server with servername and database name using pyodbc in Python: conn = pyodbc.connect('Driver={SQL Server};Server=YOUR_SQL_SERVERNAME\SQLEXPRESS;Database=YOUR_DATABASE_NAME;Trusted_Connection=yes;') pyodbc(Open Database Connectivity) to Connect a MySql Server in Python: //a href="https://docs.devart.com/odbc/mysql/python.htm"> How to Connect to MySql Server with pyodbc in Python //a href="https://stackoverflow.com/questions/3982174/pyodbc-and-mysql"> Stackoverflow site for pyodbc: Examples to Connect to MySQL pyodbc(Open Database Connectivity) to Connect MS SQL Server in Python: //a href="https://docs.devart.com/odbc/sqlserver/python.htm"> How to Connect to MS SQL Server in Python //a href="https://datatofish.com/how-to-connect-python-to-sql-server-using-pyodbc/"> How to Connect to MS SQL Server in Python //a href="https://stackoverflow.com/questions/33725862/connecting-to-microsoft-sql-server-using-python"> pyodbc: Examples of ODBC in Python //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-1-configure-development-environment-for-pyodbc-python-development?view=sql-server-ver15"> Installation pyodbc to Connect to MS SQL Server in Python //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-3-proof-of-concept-connecting-to-sql-using-pyodbc?view=sql-server-ver15"> Stackoverflow site for pyodbc to Connect to MS SQL Server in Python //a href="http://www.sqlines.com/oracle/datatypes/clob"> Column Data Types to store a Large Text data to Create a Table with in MS SQL Server or other database server How to Create a Table in a SQL Server from CSV/TSV Text files //a href="https://docs.microsoft.com/en-us/sql/t-sql/statements/bulk-insert-transact-sql?view=sql-server-2017"> How to Create a Table from a File with Bulk Insert with MS SQL Server //a href="https://stackoverflow.com/questions/14330314/bulk-insert-in-mysql"> How to Create a Table from a file with Bulk Insert in MySQL Server //a href="FullBulkImportAWIDSQLServer.sql"> Example of a Script to Create Multiple Tables from different files with Bulk Insert in MS SQL Server Earlier Features to Handle Big Data in Relational Database Server Advanced Data Types: BLOB (Binary Large Object) or Text/CLOB (Character Large Object) in MySQL or MS SQL Server How to Create and Insert to a Table with a Column of Large Text Data or Image Data in a Relational Database Server: //a href="https://www.mysqltutorial.org/mysql-text/"> Text Data Type in MySQL //a href="https://www.tutorialspoint.com/What-is-TEXT-data-type-in-MySQL"> What is TEXT data type in MySQL //a href="https://www.tutorialspoint.com/what-is-the-difference-between-blob-and-clob-datatypes "> Difference between blob and clob datatypes //a href="https://dev.mysql.com/doc/refman/8.0/en/blob.html "> Blob Data type in MySQL //a href="https://stackoverflow.com/questions/10729824/how-to-insert-blob-and-clob-files-in-mysql"> how to insert blob and clob from files in mysql //a href="https://stackoverflow.com/questions/752277/cannot-insert-string-into-mysql-text-column "> how to insert text column in mysql For Those Who Want to Use CLR Table Function to Create a Table in SQL Server -- This is NOT For CIS492/593 //a href="HowtoSETUPUDTASPNET.pdf"> How to Set Up ASP.NET with SQL Server //a href="HowToDebugSQCLRNOExamples.pdf"> How to Debug CLR UDF, CLR UDT, CLR TVF Lab 2 on JSON Data Processing Important Notes and Corrections: I clearly mentioned and covered how to map from Semi-structured data JSON to relational database schema in class. Coverting all to one big messy table is NOT a good solution for Lab2. Using an bad API that converts from a JSON file to a CSV file is NOT a solution of Lab2. Note that you need to transform business.json file only for Lab2, not all of 5 json files from the Yelp site. Yelp Business Data Set for Lab2: Yelp data set and documentation for JSON File Structures You can directly download the most recent data sets from Yelp site: Yelp Full Data Set Also Avaliable in the Big data Lab below: //a href="https://eecs.csuohio.edu/~sschung/cis612/YelpDataSets.zip"> Zip file for 5 JSON files from Yelp data Challenge 2017 (Or Download directly from the Yelp site below for 2020 data sets!) Note: You Have to Use 7-Zip to unzip the zip file. Some other compression software might not be able to unzip correctly. The full JSON data files from the Yelp site might have a few incorrect JSON syntax detected in the data file or invalid line feed, which is common. Correct them before processing if detected. See FAQs for how to correct invalid JSON data Invalid JSON format handling: If there is any invalid data format is found in the input file, you can change it to the correct JSON format. For example, $$ in the “Price Range” key value pair in your input json file, the value $$ is not in quotes and this will cause to fail. Suggested Solutions: Replace $$ with 2 (meaning the price level is 2 in scale 1 - 5) in the file and try to parse the corrected file in your program. If , is missing between objects, add it. The given OneBusiness.json file and business100.json file in JSONData.zip are the corrected files. Lab 3 on Social Media Logging Data to NoSQL Database MongoDB IMPORTANT NOTES for LAB3 !!! Due to the Recent Changes in Twitter - They Charge for Data Collection, You Don't Need to Collect the New Streams of the Twitter Logging Data for Lab3 You Can Collect Any Social Media Messages (in the JSON Logging Structure) Such as FaceBook, Reddit for Lab3 ! You Can Also Use Old Twitter Data You Collected Before for Lab3 If You Can't Collect Any Social Media Logging Data, Do Lab3_1 with Yelp Data Set Below as Lab3 //a href="TwitterLoggingStructureJSON.pdf">Twitter Logging Structure in JSON Twitter API 1.1 is depreciated and wont be available for new developers: //a href="https://twittercommunity.com/t/deprecation-announcement-removing-compliance-messages-from-statuses-filter-and-retiring-statuses-sample-from-the-twitter-api-v1-1/170500"> Twitter API 1.1 is depreciated The New Structure for the 2.0 API for tweets: //a href="https://developer.twitter.com/en/docs/twitter-api/data-dictionary/object-model/tweet"> the New Structure for the 2.0 API for tweets //a href="FAQs_Lab3_3_TwitterDataCollection.pdf"> FAQs to Get Twitter Developer's Account Note That the Recent Changes of the Twitter Site Seem to Affect Their Response Time to Process Applications for Developer's Account. Due to the Delay on the Twitter Site Response Time, You Need to Apply ASAP to Collect the Twitter Streaming Data in Time Twitter Stream Data Collection Collect at least 10,00 Tweets Talking about either One of the Following Topics of Your Choice. Add More Related Keywords As Needed to Your Chosen Topic to Collect As Many Related Tweets Possible: Suggested Topics: 1. Any New Major Movie or Product That Was Released Recently if Any (For example, IPhone - IPhone 12, IPhone Mini) Or 2. Any Major News (For example, Russian Invasion to Ukraine) Or 3. President or Any Person of Interest, or Any Two Candidates in an Election Or 4. Covid, Corona Virus, Covid-19 Or 5. Any Topics of Your Interest as long as there are big enough to collect more than 20,000 Tweets Twitter Data Collection Setting Up: (The Twitter Data Collected will be used for Lab3 and Can Be Used For Your Final Project later) NOTE that to collect the Twitter stream data in real time, you need to apply for their developer’s account in the Twitter Deveoper's site and get a permission to get credentials for a token and keys. This process usually takes 3-4 days or one week for Twitter to respond. You HAVE TO start your application ASAP. Don’t wait until the last day. See the example project and sample codes for the step by step procedures for this. Read everything posted in the links below. Apply a Twitter Developer's Account ASAP for Twitter Stream Data Collection ! Start ASAP Since it Will Take a Week to Obtain a Twitter Developer's Account !! For Your Twitter Developer's Account Application, Choose the most Common Account type. Do not choose an Academic Reserach Account (You Will Be Asked More Questions). Do NOT Blindly Copy Those Sample Answers in the Class Webpage for Your Answers ! Rephrase/Modify Them in Your Words For Your Case. //a href="Example of Answers to Obtain Tweeter Developer.pdf"> See Examples of the Answers for the Questions from Twitter to Get a Twitter Developer's Account //a href="SampleAnswersForStudent_TwitterDeveloperAccountApplication.pdf"> (This is For an Academic Research Account, which You Don't Need to Apply) See Sample Answers for the New Questions from Twitter to Apply Twitter Developer's Account //a href="FAQs_Lab3_3_TwitterDataCollection.pdf"> FAQs to Get Twitter Developer's Account For Your Twitter Account Application For the Project Site if Asked, Provide Your Lab3 Specification above. you Can also Provide the CIS612 Project Site and Research Project Description for Big Data and Data Scientist //a href="https://eecs.csuohio.edu/~sschung/cis612/CIS612Project.html"> CIS 612 Project Site to Provide in Your Application for Twitter Developer's Account //a href="Faculty_Led_Project_Description__Social MeadiaSentimentAnalysisSystem_SunnieChung.pdf"> To Answer with Sample Research Project Description for Big Data and Data Scientist Twitter API 1.1 is depreciated and wont be available for new developers: //a href="https://twittercommunity.com/t/deprecation-announcement-removing-compliance-messages-from-statuses-filter-and-retiring-statuses-sample-from-the-twitter-api-v1-1/170500"> Twitter API 1.1 is depreciated The New Structure for the 2.0 API for tweets: //a href="https://developer.twitter.com/en/docs/twitter-api/data-dictionary/object-model/tweet"> the New Structure for the 2.0 API for tweets How to Collect a Twitter Stream to a JSON File then Insert to MongoDB //a href="TwitterLoggingStructureJSON.pdf">Twitter Logging Structure in JSON Note some Tweets Don't have the Retweet part of info. Try to Collect Retweets Together. How to Get Twitter Stream in JSON file: //a href="CollectingTweetStreamingDataintoJSONfileWithPython_Sp2022.pdf"> Step by Step Guide on How to Get Twitter Stream in JSON file in Python (After the Tweepy API Version 4.0 as of Spring 2022) NEW POST !! by TA Yixi Luo **************** //a href="Fetching tweets data from Twitter with Python.pdf"> Step by Step Guide on How to Get Twitter Stream in JSON file in Python (Before The Tweepy New Version 4.0) How to Get Twitter Stream in JSON file then Insert into MongoDB: //a href="StepByStepInstructions_CollectingTwitterStreamFromDeveloperAccount.pdf"> Tutorial on How to Get Twitter Stream to Insert into MongoDB in Python As of 2021 Before the New Tweepy Version 4.0 Note: This tutorial uses last year (2021)'s Tweepy Streaming Library. It has been upgraded to a new version as of early 2022. See The Tutorial for the New Version of Tweepy Streaming Lib Above Blindly copy and paste of the codes in this tutorial wouldn't work because of the new version of Tweepy Streaming Lib //a href="FetchingTweetsfromTwitteronHDFSMongoDB.pdf"> Other Related Documentations and Examples from Twitter sites and Python for Twitter If You Failed to Get a Twitter Developer's Account, Use this Twitter Raw data set in this site for This Lab //a href="https://www.kaggle.com/code/prathamsharma123/clean-raw-json-tweets-data"> Kaggle: Clean Raw JSON Tweets Data site //a href="RawJsonTwitterData.zip"> RawJsonTwitterData.zip //a href="coronatweets_11_53.zip"> RawJsonTwitterData: coronatweets_11_53.zip If You Can't Collect Any Social Media Logging Data, Do Lab3_1 with Yelp Data Set Below as Lab3 The Recent Changes of the Twitter Site Seem to Affect Data Collection of Developer's Account. Due to Policy Changes on the Twitter Site, You Can Do Lab3_1 with Yelp Data Instead of Lab3 on Social Media Logging(Stream) Data Lab3_1 with Yelp Data Set: Semi-Structured Data Processing with Semi-Structured Database Server MongoDB: 1. Creating a Semi-Structured Database - MongoDB Collections in a MongoDB Server for business.json and review.json files from the Yelp site 2. Writing Aggregation Pipelinig //a href="CIS593_Lab3_Yelp_MongoDB_BusinessDataOnly_ExtraCredit_Join.pdf"> Lab3_1 on Writing Aggregation Pipelinig on Mongo DB Collections from json data files from Yelp site Make Sure to Use All the Documents in business.json to create a Collection named business for this Lab Data Sets for Lab3_1: //a href="YelpDataSets.zip"> Zip file for 5 JSON files from Yelp data Challenge 2017 (They may not be in valid JSON format) or Download directly from the Yelp site below for 2020 data sets! Note: If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip. You can directly download the most recent data sets from Yelp site: //a href="https://www.yelp.com/dataset"> Yelp Site for Data Sets //a href="JSONData.zip"> //a href="Lab3JSON.zip"> Zip file for One business data and 100 busineess from Lab2 (in valid JSON format) If you have invalid JSON file problem, the JSON file might need to be cor rected through the JSON validator to detect the errors first then correct the JSON syntax to be imported. The data files in the Zip file are all coming from the yelp site. They sometime have the incorrect JSON syntax, which need to be corrected to make it work. //a href="FAQOnLab3_YelpMongoDB.pdf">FAQs On Lab3_1 on Yelp Data with MongoDB See the example of how to correct an invalid json file below. Scroll down for the part. //a href="CIS593_Lab2_JSON_FAQs.pdf"> FAQs for JSON Processing in General from Lab2 FAQs More FAQs on Lab3 on MongoDB: Q: Can we use the 2019 dataset (JSON) from yelp website for our Lab 3_1? Or we should only use 2017 JSON data? And Do we need to create CSV for each query as well for the count result? A: Either data set is ok. Some people have a Zip error with the 2017 data set. If then use the 2019 data set from the Yelp site. Conversion to CSV file is not required for This Lab3_1. Lab3 Mongo DB Guides: Mongo DB Setup: MongoDB Installation: //a href="CIS593_612_MongoDB Setup Atlas Mongosh.pdf"> Step by Step MongoDB Installation with Atlas and Client Mongosh(by TA Sowmya Chinthalapudi) //a href="https://docs.mongodb.com/manual/installation/"> MongoDB Download and Installation //a href="https://docs.mongodb.com/manual/tutorial/install-mongodb-enterprise-on-windows/"> MongoDB Installation Guides //a href="https://www.mongodb.com/download-center#production"> Full Menu for MongoDB Download //a href="https://www.mongodb.com/download-center#enterprise"> MongoDB Enterprise Download //a href="https://www.mongodb.com/download-center#compass"> MongoDB GUI Client Compass Download //a href="https://docs.mongodb.com/manual/"> Full Menu of MongoDB Manuals for Getting Started //a href="https://docs.mongodb.com/mongodb-shell/#mongodb-binary-bin.mongosh"> MongoDB Cilent: Mongo Shell (mongosh) to start //a href="https://docs.mongodb.com/mongodb-shell/connect/#std-label-mdb-shell-connect"> MongoDB Shell to Connect //a href="https://docs.mongodb.com/manual/tutorial/write-scripts-for-the-mongo-shell/"> MongoDB Cilent: Writing Scripts for Mongo Shell //a href="https://docs.mongodb.com/drivers/pymongo/"> How to Install pymongo driver and Connect to MongoDB Server in Python application as a client See MongoDB Lecture Notes Section for MongoDB CRUD Queries and more details //a href="https://docs.mongodb.com/guides/server/import/"> How to Import Data file to MongoDB //a href="https://docs.mongodb.com/manual/reference/program/mongoimport/"> Mongo import //a href="MongoDBQueryExamples.pdf"> Sample Runs of Mongo DB Queries //a href="https://www.mongodb.com/compatibility/json-to-mongodb"> How to Import a json file to MongoDB //a href="https://www.mongodb.com/languages/python"> How to Use MongoDB in Python for CRUDE: DB/Collection Creation, Insert, and Querying //a href="https://www.geeksforgeeks.org/how-to-import-json-file-in-mongodb-using-python/"> How to Import a JSON file to MongoDB in Python //a href="https://stackoverflow.com/questions/70089317/how-to-do-a-word-count-in-mongodb"> Example of MongoDB Aggregation Pipeling for Word Count //a href="https://stackoverflow.com/questions/42941682/storing-json-data-into-a-variable-using-python-when-inserting-into-mongodb"> How to Save MongoDB Query Results into a variable //a href="https://stackoverflow.com/questions/20769621/saving-the-result-of-a-mongodb-query"> How to Save MongoDB Query Results Lab 4: Text Analytics (Text Mining) with Information Retrieval and Natural Language Processing Methods Lab 4 (Lab4_3) on Information Retrieval Methods for Content Based Document Search Engine IMPORTANT NOTE !! - Note that Building NLP Data Pipelining with Lemmarizer/Stemming, POS and NER is REQUIRED for Lab4_3 ! (It is NOT an Extra Credit) - You have to build Inverted Index over ALL the State Union Address Text documents ! IMPORTANT NOTES on LAB4 Part 1 (Phase 1): The Lab4 of the Big Data course REQUIRES to build: Part 1 (Phase 1): 1) an Inverted Index on all the ~200 Union Address texts 2) with building a NLP data procesing pipelining with 3 required NLP techniques - Lemmatizer, POS, NER in a correct sequence while parsing. Part 2 (Phase 2): 3) Build a TF-IDF weight matrix for 10 given addresses to calculate Cosine Similarity between them to identify the closest document to a user given topic word set (as a Query document) to answer The focus of Lab4 Par 1 (Phase 1) in the Big Data class that requires: 1. Unstructured Text Document Processing with NLP Data Pipelining to Extract Information and, 2. To Build and Manage the Extracted Information from a Large Collection of Documents in an Inverted Index for Retrival in the Next Phase of AI for Similarity Matrix calculation in Real-time for a Given User Question in a Database Server either a SQL Server or MongoDB Server to Retrive the TF values with each Doc_ID and DF value for each given user Topic Word in a query Input Files for Lab4_3 Part 2 (Phase 2): //a href="https://eecs.csuohio.edu/~sschung/CIS593/Text_Mining_ConstructingInvertedIndex.pdf"> Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example) You Can Choose to Extend One of the Labs (Lab4_2 or Lab4-3) Combined with the Inverted Index on a Big Collection of Documents as a Final Group Project ! Lab 4 (Lab4_3) Part 1 (Phase 1): - Building an Simplified Inverted Index in a SQL Server on the Given Full State Union Addresses (Required for All for Lab4) - Due By April 11 For Inverted Index to Build, You Can Simplify to One Table with (Term, Doc#, TermFreq) for Part 1 of Lab4_3 Part 2 (Phase 2): The Next Phase for AI to Calculate Document Similarity to a User Given Query (as a Document) in Weight Matrix in TF-IDF Scoring function to Return the most similar document to a User Query document as an Answer in Real time: Required for CIS492/DSA469 for Lab4_Part2: - Building a Similarity Matrix for a Pair of Question and Document and each of the Given 10 State Union Addresses Or - Building a Similarity Matrix for Every Pair of two Documents for the Given 10 State Union Addresses Required for CIS593 for Lab4 Part2: - Building a Similarity Matrix for a Pair of Question document for the Given Entire State Union Addresses Or - Building a Similarity Matrix for Every Pair of two Documents for the Given Entire State Union Addresses 1. Building a Full Inverted Index on the Entire Set of State Union Addresses in a SQL Server is an Extra Credit (50%) for CIS492/DSA469 as in the Lecture Note as below: Inverted Index Scheme as in the Lecture Note: 1)Dictionary table (Term, TotalDocsFreq, TotalCollectionFreq) and 2)Posting Table(Term, Doc#, Term_Freq) 2. Building a Full Inverted Index in JSON Semi-Structure on the Entire Set of State Union Addresses in a MongoDB Server is an Extra Credit (75%) for CIS593 You Have to Design Semi-Structed Inverted Index in JSON to Store all the columns in Dictionary Table and Posting Table in one JSON document per one union address 3. Building an Full Inverted Index on the Entire Set of State Union Addresses with NLP Pipelining for each Sentence to Extract Context Aware Information with POS and NER Taggers along with the basic required NLP processing and Store them either in a SQL Server (ex: MySQL) or MongoDB Server for Real-Time Retrieval (75%) in the Next Phase for AI to Calculate Document Similarity to a User Given Query (as a Document) in Matrix in TFIDF Scoring function to Return the Answer in Real time. Text Preprocessing Library in Python SpaCy: //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing Lab 5 on Classification with Machine Learning: Requirements for Lab5: 1. Experiment Your Classification with two Different MLs: 1_1. One for Probablistic Based ML: Decision Tree, Random Forest, or Naive Bayes 2_2. One for Numerical Approach ML: ANN, or SVM 2. For Each ML, Find the Best Hyperparameters (input parameters) to Find the Best Performing model For Example, with ANN, Repeat training to find the optimal hyperparameters (Learning rate, the Numer of Hidden Layer) for the best fit(model) for the given data (150% If You Do Lab5_1 Instead of Lab5) Implement ANN Classification with //a href="breast_cancer_dataset.csv"> Health Data Set: Patient Data for Breast Cancer Prediction for Binary Classification For Lab 5, You Can Choose to Use Any of the Following Data Sets: Note that the Data files and Data Description Files are text files, you can open them as a text file with any text editors like Notepad or Notepad++ 1. Adult Profile Data Set to Predict Income Level < 50k or NOT from Adult Data set //a href="http://archive.ics.uci.edu/ml/datasets/Adult"> UCI Site for Census Adult Profile Data Set and Data Description for Binary Classification 2. Wine Data Set with either DT or NN to Predict Wine Quality in Scale 1 - 10 //a href="https://archive.ics.uci.edu/ml/datasets/Wine+Quality"> UCI Site for Wine Data Set and Data Description for Multi-Class Classification 3. Patient Data Set for Breast Cancer Prediction for Binary Classification to Predict whether the Patient has a Breast Cancer or not ML Algorithms -- ANN or SVM Reqire Data Preprocessing: Tutorial for Data Preprocessing Normalization and One Hot Encoding for ANN or SVM //a href="https://scikit-learn.org/stable/modules/preprocessing.html#preprocessing-categorical-features"> Categorical Data transformation Methods with Binarization (One Hot Encoding) //a href="https://visualstudiomagazine.com/articles/2013/07/01/neural-network-data-normalization-and-encoding.aspx"> Data Preprocessing Methods for ANN or SVM //a href="https://www.mltut.com/implementation-of-artificial-neural-network-in-python/"> Code Walk Through with a Simple Example using Panda for Classification with Neural Network ******************* You Can Choose Your Own Data Set to Extend Lab5 for Final Group Project ! Data Set Repository for Classification: //a href="https://www.kaggle.com/datasets"> Kaggle Data Set Repository //a href="http://archive.ics.uci.edu/ml/datasets.php"> UCI Data Set Repository for Classification //a href="https://healthdata.gov/search/type/dataset"> Health Data Set Repository for Classification //a href="https://dev.socrata.com/"> Socrata Open Data API Extra Credt Lab 5_1 on Sentiment Analysis with Machine Learning for Classification (Not Required for Everyone): Extra Lab 5_1 on Sentiment Analysis with Machine Learning for Classification: Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set Classification Goal: 1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review. 2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star) Useful Big Data Analytic Tools Choose your System/Tool/Platform to Set Up and Get Used to: Python Analytics Tools and Tutorials : Machine Learning : Other Machine Learning Platforms: Basic R Tutorials : Extra Credit Lab 6 on MapReduce and Hadoop (HDFS): Extra Credit Lab (Not Required for Everyone) Hadoop Set Up Instruction Sites: //a href="http://hadoop.apache.org/docs/r1.0.4/single_node_setup.html"> Hadoop single node setup //a href="http://hadoop.apache.org/docs/r1.0.4/cluster_setup.html"> Hadoop Cluster setup //a href="http://hadoop.apache.org/docs/r1.0.4/mapred_tutorial.html"> Map Reduce Tutorial on Hadoop Lab Guides: Newest on the Top //a href="Lab4_1InstallHadoopDataNodeError.pdf"> Hadoop Installation: How to Fix When Data Nodes are not Running (2018) //a href="https://stackoverflow.com/questions/11889261/datanode-process-not-running-in-hadoop"> Help Site on How to Fix When Data Node are not Running (2018) //a href="HowtoExecuteMapReduceinEclipse.pdf"> Procedure to How to Execute MapReduce in Eclipse to Run a Wordcount Job (2018) //a href="Lab_4_1_Alex_Chengelis.pdf"> Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 (2017) //a href="CIS 612_Lab4_1_HadoopSetUpHalley.pdf"> Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 - Permission issue (2017) //a href="Lab4_1_CIS612_HDFSInstall_MacSonal.pdf"> Procedure to Install Hadoop and Run a Wordcount Job on Mac (2017) //a href="CIS_612_Lab_4_1HadoopWordCount_Danielle"> Guideline4 for Passwordless SSH for Lab4_1 (2016) //a href="HadoopWordCountJamesTench_2015.pdf"> Guideline2 for Lab4_1 (2015) //a href="GuideforBigDataProcessingLab4ByRyanChesla.pdf"> Guideline1 for Lab4_1 (2014) A good Instruction Site for Installing and Running Hadoop //a href="Setting-up-Hadoop-made-easy.pdf"> Installing Hadoop If you have a trouble installing Hadoop from the above site with not seeing the name node, You need to delete the temp files created in standalone mode and reformat the namenode. For setting up a passwordless SSH, see Ganesh's Lab4 below as well. //a href="HadoopAssignment4Ganesh.pdf"> Installing Hadoop and running a wordcount job by Ganesh VAVILAPALLI //a href="https://www.youtube.com/watch?v=MoKW5eY5yVY"> Video for Installing Hadoop shared by Prashant Patel //a href="TroubleShootingTipswithHadoop.pdf"> Trouble Shooting Tips for Installing Haddop on VM //a href="Lab_MRHadoop_Um.pdf"> Guideline3 for Lab4_1 on Mac (2013) For Those who Want to Set Up Your HDFS Cluster On EC2 Amazon Cloud, See the Cloud Section at the end of the Class Lecture Note Section for a Student Account. Supporting Contents for Labs : How to Create a Web Application with Java Based Application Server with MS SQL Server: Amazon Cloud: //a href="https://aws.amazon.com/rds/">Amazon RDS (Relational Database Service) //a href="http://aws.amazon.com/ec2/">Amazon Elastic Cloud Computing (EC2) for Web service //a href="http://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/TUT_WebAppWithRDS.html">How to Create Amazon Virtual Hosting (EC2) with a Web Server and Amazon Database Server (RDS) Microsoft Cloud Azure: //a href="https://azure.microsoft.com/en-us/">Trial account for Microsoft AZURE Cloud //a href="MicrosoftAzureTables.pdf">How to Create/Retrieve a Table in Microsoft AZURE Cloud //a href="https://code.msdn.microsoft.com/Fix-It-app-for-Building-cdd80df4"> Sample Project on Microsoft AZURE Cloud //a href="http://www.asp.net/aspnet/overview/developing-apps-with-windows-azure/building-real-world-cloud-apps-with-windows-azure/unstructured-blob-storage">How to Create/Retrieve BLOB data in Microsoft AZURE Cloud //a href="http://www.davidchappell.com/Azure_Services_Platform_v1.1--Chappell.pdf">AZURE Cloud //a href="windows_azure_sql_database_tutorials.pdf">How to Create a SQL Database Server in Microsoft AZURE Cloud //a href="http://www.tutorialspoint.com/microsoft_azure/index.htm">Tutorial for MS AZURE Cloud Useful Resource Sites: |
| Class | Chapter / Topic / Specific Objectives / Activities |
| 1 |
|
| 2-5 |
|
| 5-6 |
|
| 10-11 |
|
| 11-12 |
|
| 12-13 |
|
| 13-14 |
|
| 15 |
|
| 6-7 |
|
| 16 |
//a href="ENACh13final-Disks-FileStructure-Hashing.pdf"> Project Presentation |
==> Completion of Homeworks/Labs is required for obtaining a passing grade.
| This
is a tentative scale and |
Letter |
Quality Points |
|
||
| A |
> 93% |
A: Outstanding (student's performance is genuinely excellent) | |||
| A- |
90% - 93% |
||||
| B+ |
87% - 90% |
||||
| B |
82% - 87% |
B: Very Good (student's performance is clearly commendable but not necessarily outstanding) | |||
|
|
B- |
80% - 82% |
|||
|
|
C |
75% - 80% |
C: Good (student's performance meets every course requirement and is acceptable; not distinguished) | ||
| D | 65%-75% | D: Below Average (student's performance fails to meet course objectives and standards) | |||
|
|
F |
<65% |
F: Failure (student's performance is unacceptable) | ||
|
ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106. |
Programming standards
|
|
|