Magnus Mårtensson
Microsoft Regional Director, Azure MVP, CEO Loftysoft
Avirag Jain
Director & CTO R Systems
Mahesh Chand
Founder C# Corner, CEO Mindcracker
Chris Gali
CEO & Co-Founder Graphite
Subinder Khurana
Chief Architect StoryProcess, Founder NASSCOM DeepTech Club
Bryan Rishforth
Investor, Chairman Graphite
Bryn Everson
Director Biz Dev Graphite
Raj Tiwari
Digital Transformation Leader, Futurist and Visionary
Joseph Guadagno
Microsoft MVP, Lead Quicken Loans
Nikita Sachdev
Entrepreneur, Blockchain Enthusiast & Advisor, Social Media Influencer
Doug Wagner
COO & Founder Adapt Technical Group
Ritesh Modi
Architect, Senior Evangelist, Cloud Architect
Crystal Wenrick
Director Communications Mindcracker
Allen O’Neill
Microsoft MVP, Consulting Engineer/Architect
Praveen Kumar
CEO MCN Solutions
Chris Love
Founder Love2Dev, Microsoft MVP, Author
Sanjay Vyas
Microsoft Regional Director, Microsoft MVP, Founder & CEO SkillLabs Technologies
Veena Sarda
Deep Learning Consultant, Author
Sekhar Srinivasan
C# Corner MVP, Microsoft Certified Trainer, Pluralsight Author
Lalit Bansal
Founder & CEO - EIY SYS
Navdeep Garg
CEO Revinfotech
Prakash Tripathi
Tech Manager/Leader, Microsoft MVP, Blogger
Bhavna Jain
Breakthrough Consultant
Naveen Sharma
Enterprise Architect, Leadership Coach, Author
Vidya Vrat Agarwal
Principal Architect, Microsoft MVP, Author
Sheetal Agarwal
Founder Clownselors, Medical Clown, Trainer
Abhishek Kant
Founder GTM Catalyst
Vishnu Saran
Founder & CEO VoiceQube
Sandeep Soni
Founder & CEO Deccansoft, Microsoft Certified Trainer
Parveen Malik
AVP InfoSec & Vulnerability Management, Information Security Expert
Nitin Pandit
Microsoft MVP, Developer Evangelist, Author
Niloshima Srivastava
C# Corner MVP, Tech Architect, Trainer, Blogger
Bala Chirtsabesan
Senior Software Engineer at Microsoft, Author
Manoj Mittal
Sr. Technical Architect, C# Corner MVP, Author
Chandni Di
Co-Founder Voice of Slum
Vithal Wadje
Technical Lead, Microsoft MVP, Author
Shivam Ahuja
Founder SkillCircle, Business Mentor
Chervine Bhiwoo
Solution Architect, Microsoft MVP, Author
Saurabh Jain
Vice President Paytm, Founder Fun2Do Labs, Author
Vinay Solanki
Head IoT at Lenovo, Founder IoT-NCR
Anshu kumari
Founder Blockchainkids, Inventor, Trainer
Amit Singal
CEO Startup Buddy
Dev Pratap
Co-Founder & CEO Voice of Slum
Amey Vartak
Technology Consultant, Full Stack Developer, C# Corner MVP, Author
Viswanatha Swamy
Principal Software Engineer, C# Corner MVP, Author
Sanket Verma
Research Engineer @ Ballistics (Forensics) and Chair, PyData Delhi
Sourabh Somani
Lead Developer, Microsoft MVP, Author
Abhishek Mishra
Software Architect, C# Corner MVP, Author
Siddharth Vaghasia
Technical Consultant, C# Corner MVP, Blogger
Bassam Alugili
Senior Software Specialist, Database Expert
S Ravi Kumar
Solution Architect, C# Corner MVP, Author
Sundaram Subramanian
Full Stack Developer, C# Corner MVP, Speaker
Deepesh Somani
Solution Architect, Microsoft MVP, Author
Debasis Saha
Technical Project Manager, C# Corner MVP, Blogger, Author
Vipul Jain
Software Architect, C# Corner MVP, Author
Akshay Patel
Technical Architect, Microsoft Certified Trainer, C# Corner MVP, Author
Stephen Simon
RPA Developer, Evangelist, Author
Vivek Sharma
Founder Kingster636, AR/VR Specialist
Jeetendra Gund
Technical Lead, C# Corner MVP, Author
Sujal Beniwal
AI Enthusiast, Student
M Viknaraj
Microsoft MVP, Azure Architect, Author
Prasham Sabadra
Software Architect, C# Corner MVP, Trainer, Author
Aakash Maurya
Senior Developer, C# Corner MVP, Speaker
Ankit Sharma
Senior Software Engineer, C# Corner MVP, Author
Mangesh Gaherwar
Team Lead, C# Corner MVP, Author
Viral Jain
Technical Consultant, C# Corner MVP, Author
Bhasker Das
Solution Architect, Evangelist
Manish Dwivedi
Associate Project Manager
Ck Nitin
Programmer, Author
Rohit Gupta
Technical Trainer, Author
Manish Tewatia
Full-stack Marketer, UX Designer
Bhavya Gaur
Technical Illustrator
Rohit Tomar
SEO/SMO Expert
Web Track
Cloud & Data Track
Dev Track
Registration & Breakfast
Future of Desktop Apps with JS (ElectronJs)
Nitin Pandit
Building Serverless Microservices Using Microsoft Azure
Vithal Wadje
Innovating RPA: A Robot for Every Person
Stephen Simon
Managing Cloud Storage Accounts using Logic Apps
Viknaraj Manogararajah
Data visualization using Python
Sekhar Srinivasan
Going Cross platform with AR Foundation
Vivek Sharma
Keynote
Managing your Azure dependencies in ASP.NET Core apps using VS
Bala Chirtsabesan
Securing Applications on Intelligent Azure
Abhishek Mishra
Getting started with Blazor the Framework of Future
S Ravi Kumar
Lunch
Build Progressive Web Apps using Angular 9
Debasis Saha
Build and deploy to any platform using Azure DevOps
Chervine Bhiwoo
Deep Dive in Azure Service Bus
Akshay Patel
Build a Native Mobile Application using React Native and JavaScript
Joseph Guadagno
Making sense of Web Job, Web Job SDK and Functions in Azure
Prakash Tripathi
CloudFront Distribution in AWS
Viral Jain
Tea Break
Introduction to PowerBI
Aakash Maurya
Build Advanced SPFx solutions with React and Graph API
Siddharth Vaghasia
Build Business Intelligence Analyst (BIA) Skills
Sundaram Subramanian
Deep dive of Power Platform – AI BUILDER
Prasham Sabadra
Panel 1
What's new in SharePoint development
Vipul Jain
Build a SSO (Single Sign On) based Native JavaScript application with Microsoft Identity within 10 minutes
Manoj Mittal
Panel 2
Applications and working of AI
Veena Sarda
Deploying serverless API's with .Net core 3.0 on AWS & Azure
Amey Vartak
Panel 3
Blockchain with .NET Core (Ark)
Anshu Kumari
Closing Note & Prize Distribution
Dev Track
Cloud Track
Architecture Track
Emerging Tech Track
Registration & Breakfast
Creating Full-Stack Web Apps Using Server-Side Blazor
Ankit Sharma
Real time face recognition with MS Cognitive Services
Niloshima Srivastava
Building Scalable APIs with GraphQL
Jeetendra Gund
Future of development with AI and Blockchain
Navdeep Garg
Debugging Tips and Tricks with Visual Studio 2019
Joseph Guadagno
Azure Containers
Abhishek Kant
Enterprise Architecture
Naveen Sharma
Bot Framework - learn it fast and look like a boss!
Allen O’Neill
Keynote
.Net Core & C# 8 Performance
David McCarter
Working with Azure kubernetes services
Ritesh Modi
Becoming an Architect
Vidyavrat Agarwal
Why Techies Need to Learn Product Management
Saurabh Jain
Lunch
Build a rules engine in .Net Core
Sanjay Vyas
Building CI and CD Pipeline using Azure DevOps
Sandeep Soni
Entity Framework Core - Tips and Tricks, Performance Optimization, and Tuning
Bassam Alugili
Hacking your way into Data Science
Sanket Verma
Speed up your .Net Core Website
Sourabh Somani
Azure
Magnus Mårtensson
Demystifying Open Distro for Elasticsearch
Suman Debnath
Future of Data
Shivam Ahuja
Tea Break
gRPC with C# and .Net Core
Mangesh Gaherwar
Panel 1
Essentials of Cloud security
Parveen Malik
Power platform and Dynamics 365
Deepesh Somani
Microservices - the gRPC Way
Viswanatha Swamy
Panel 2
Reserved
Reserved
Closing Note & Prize Distribution
Scraping the Web with HtmlAgilityPack in C#
The internet is a sprawling archive of product listings, job adverts, weather records, and government open data, and a surprising amount of it can be turned into structured information with a few lines of C#. For developers in Brisbane or Adelaide who want to pull prices from real estate portals, harvest listings from Australian job boards, or aggregate news from local council sites, HtmlAgilityPack remains one of the friendliest .NET libraries for the job. It wraps the messy realities of HTML into an XPath-queryable DOM, which means you spend less time fighting malformed tags and more time getting useful data out of the page.
Behind every reliable scraper sits a small amount of preparation. Before writing a single line of code, it helps to inspect the target page in a browser, look at the network requests it makes, and decide whether the data you need is rendered server-side or arrives later through JavaScript. HtmlAgilityPack works best when the HTML arrives in a single response, so static content, server-rendered markup, and well-behaved APIs are all good candidates. A quick visit to a site like realestate.com.au or seek.com.au, followed by a careful look at the developer tools, will usually tell you whether HtmlAgilityPack is the right tool or whether you should reach for something heavier like Playwright.
The library itself is distributed as a NuGet package, which means a fresh console project on a developer's machine in Perth or Hobart can be scraping in minutes. Once installed, you load a page from disk or from an HttpClient response, parse it into an HtmlDocument, and start querying. The rest of this walkthrough builds a small, real-world scraper step by step, from the initial project setup right through to exporting clean data and respecting the legal boundaries that apply in Australia.
Setting up your C# scraping project
A clean starting point is a .NET 8 console application created with dotnet new console. Open a terminal, scaffold the project, and add HtmlAgilityPack with the standard dotnet add package HtmlAgilityPack command. The package restores quickly, even on a modest NBN connection in a regional town like Ballarat or Bunbury, and it has no external native dependencies, so the same binary will run on a developer's Windows laptop and on a Linux build server without modification.
Inside Program.cs, the first useful snippet is a simple async loader that fetches a page and hands back an HtmlDocument. HttpClient is injected or instantiated once and reused, which is friendlier to both the target site and to your local network. A short delay between requests, often around one or two seconds, mimics a polite human visitor and reduces the chance of being throttled. Many Australian hosts, including the ABC and the major banks, monitor request frequency closely and will serve a captcha or a hard block to anything that looks like a runaway bot.
using HtmlAgilityPack;
using System.Net.Http;
var http = new HttpClient { DefaultRequestHeaders = { UserAgent = { new System.Net.Http.Headers.ProductInfoHeaderValue("Mozilla", "5.0") } } };
var html = await http.GetStringAsync("https://example.com/listings");
var doc = new HtmlDocument();
doc.LoadHtml(html);
This pattern is the foundation of every scraper you will build with the library. From here, the real work is identifying which nodes contain the data you care about.
Navigating HTML with XPath and CSS selectors
Once the document is loaded, the next step is to describe where the interesting data lives. HtmlAgilityPack supports both XPath expressions and a LINQ-friendly selection API, which is handy when you are migrating snippets between C# and other languages. XPath is verbose but extremely expressive: you can match a <div> whose class contains the word price and grab the text of its first <span> child in a single expression.
For Australian e-commerce sites, a common pattern is to look for a container element with a stable class or data attribute, then walk down to the price, title, and image. Real estate listings on domain.com.au, for instance, nest the address, the number of bedrooms, and the price inside a repeating article tag, which makes them a textbook case for a single XPath that returns a node set. Selecting all matches is as easy as calling SelectNodes, then iterating and pulling the inner text or an attribute like href.
CSS-style selectors are available through doc.QuerySelector and doc.QuerySelectorAll, which feel more familiar to anyone who has written JavaScript. They are often shorter and easier to read, especially for deeply nested layouts. A pragmatic approach is to start with CSS selectors while exploring the page, then switch to XPath once a query needs a condition like "the second <td> in a row whose first <td> contains the word 'Sydney'". This combination covers almost every page you will encounter, from government data tables on data.gov.au to weather summaries published by the Bureau of Meteorology.
Dealing with pagination, cookies, and dynamic pages
Most useful collections live across multiple pages, and a scraper that only reads the first response is rarely useful in practice. The trick with HtmlAgilityPack is to look for the "next" link in the pagination control, follow it, and repeat until the link disappears. Storing a queue of pending URLs in a simple Queue<string> keeps the logic transparent, and a small HashSet<string> visited-set prevents you from circling back to the same listing more than once.
Cookies and redirects can complicate the picture, particularly when a site uses a session cookie to track browsing behaviour. HttpClient's CookieContainer, exposed through HttpClientHandler, handles the round-trip automatically, which is useful when scraping auction results from sites like pickles.com.au or ticket listings from Ticketek. If a site expects a referrer header, set it explicitly before each request, and if it returns a 302, follow the chain by default rather than re-issuing the original URL.
Dynamic content loaded by JavaScript is the one place where HtmlAgilityPack alone falls short. When the data you need only appears after a script runs, you have two sensible choices: use a headless browser like Playwright to render the page first, then feed the resulting HTML into HtmlAgilityPack, or call the underlying JSON API directly using HttpClient. The second option is faster and lighter on memory, and many large Australian sites, including the major banks and airlines, expose internal endpoints that return far cleaner data than the rendered DOM.
Exporting data to CSV, JSON, or SQL Server
Raw HTML is rarely the end of the journey. Once you have pulled a few thousand records, you will want to persist them in a format that downstream tools can read. CSV is the universal handshake: a row per listing, a column per field, and a header line at the top. C#'s built-in string.Join plus a small helper that escapes quotes and commas is usually enough to write a clean file without pulling in a third-party library.
JSON is a natural fit when the next consumer is a web front-end or a JavaScript-based dashboard. Records can be shaped into record types or anonymous objects and serialised with System.Text.Json, which keeps allocations low. For larger jobs, especially anything that runs on a schedule or feeds a data warehouse, a SQL Server table is hard to beat. A simple SqlBulkCopy operation can load tens of thousands of rows in seconds, which matters when you are aggregating listings from every capital city or scraping daily rainfall totals for hundreds of Bureau of Meteorology stations.
Below is a quick comparison of the most common storage choices for a small scraping project in C#. Each has trade-offs around speed, queryability, and how much extra code you have to write.
| Format | Best for | Typical size | Query support | Extra dependencies |
|---|---|---|---|---|
| CSV | Sharing with Excel, one-off analysis | Small to medium | Manual or pandas | None |
| JSON | Web APIs, JavaScript front-ends, nested data | Medium | Limited, document-style | System.Text.Json |
| SQLite | Local prototyping, single-user apps | Large | Full SQL | Microsoft.Data.Sqlite |
| SQL Server | Production jobs, shared data warehouse | Large | Full SQL | SqlClient or EF Core |
| Parquet | Analytics, long-term archival | Compact | Columnar, fast scans | Parquet.Net |
Choosing the right destination depends less on the scraper and more on who will read the data afterwards. A weekend project that catalogues coffee roasters in Melbourne might live happily in a CSV, while a nightly job that updates a national product catalogue belongs in SQL Server or a cloud database.
Staying on the right side of Australian law
Web scraping sits in a grey area, and the rules vary by jurisdiction, so it pays to be deliberate about what you collect and how you use it. The Privacy Act 1988 and the Australian Privacy Principles apply whenever personal information is involved, which means scraping names, email addresses, or phone numbers without a lawful basis can land you in trouble even if the data is publicly visible. The Office of the Australian Information Commissioner has been clear that visibility on a public web page does not, by itself, give a third party permission to collect that information for unrelated purposes.
The Copyright Act 1968 protects the expressive content of a site, so republishing full article text or product descriptions verbatim can infringe copyright even when the underlying facts are not protected. A common and usually safe approach is to store only the structured fields you actually need, such as a title, a price, and a URL, and to link back to the original source rather than copying the body content. Respecting robots.txt and the site's terms of service is also sensible, and courts in Australia have shown a willingness to enforce contractual restrictions against scrapers who ignore them.
Finally, the Notifiable Data Breaches scheme means that if you scrape personal information and then suffer a security incident, you may have an obligation to notify affected individuals and the OAIC. Encrypting scraped data at rest, limiting retention, and documenting your collection purposes are all reasonable steps that reduce both legal and reputational risk. Developers attending events like the C# Corner Annual Conference often hear seasoned practitioners describe scraping as a privilege that depends on acting responsibly, and that framing holds up well under Australian law.
1, CBD, Maharaj Surajmal Road, Near Yamuna Sports Complex, Delhi, 110032
GENERAL QUERIES
Manish Tewatia
manish@csharpcon.com
+91-9718-431-042
TICKET QUERIES
Atul Gupta
conference@csharpcon.com
+91-9910-125-804