Relevance Drives Rigor: Exploring Open Data with R
Civic Tech Chat | 2026-04-29 | 40:53
Christian Martinez shares his journey from cognitive neuroscience to civic tech, highlighting the importance of open data, reproducible research, and building accessible tools like R packages. Discover how relevance drives rigor in education and data projects, and learn practical tips for getting started with open data and open source development.
Resources and Shoutouts
NYC Open Data Portal: https://opendata.cityofnewyork.us/
NYS Open Data Portal (MTA data also lives here): https://data.ny.gov/
NYC Open Data Student Gallery Book: https://martinezc1-nyc-open-data-student-gallery.share.connect.posit.cloud/
Music Credit: Tumbleweeds by Monkey Warhol
Top Keywords
- york 0.028
- open data 0.026
- york city 0.025
- city open 0.022
- data 0.020
- city 0.019
- open 0.015
- package 0.011
- students 0.008
- york state 0.008
- data portal 0.008
- portal 0.008
Transcript
Speaker 0
0:00 – 0:52
Hello. I'm Ryan Cook, and this is Civic Tech Chat, a show that looks at the way technology, politics, and policy impacts the world around us. The tools we use, the way services are delivered, and how we talk about and set policy all shape our society. We'll gather around and have a chat about these things together and more. Before we get started, I do wanna let you all know that we've started a Discord for the podcast. There will be a link with an invite down in the episode description. Do feel free to go check that out. It's a small community right now, but hoping to grow it. It's a great way to reach out to me and let me know things that you might want us to cover or to just hang out and talk about civic tech. Christian, thank you so much for joining us here on Civic Tech Chat. Could you introduce yourself and tell us a little bit about what you do?
Speaker 1
0:52 – 1:14
Yes. Absolutely. Thanks so much for having me. My name is Christian Martinez. I am born and raised in New York. I have a master's degree in cognitive neuroscience from the CUNY Graduate Center. I currently operate my own analytics company called Angles Analytics, And a lot of the fun work that I'm doing recently is through Brooklyn College, another CUNY school,
Speaker 0
1:15 – 1:23
where I'm an adjunct. And what would you say is your personal why? A thing that drives you to get out of bed each morning and do all those things. Great question.
Speaker 1
1:23 – 1:44
A lot of it is I I sometimes describe myself as a creative opportunist where I see these opportunities to help whether it's individual people or communities, and I try to see if I can solve problems. Through those problems being solved, hopefully, I can bring people together and build some sense of community. Your background includes teaching
Speaker 0
1:44 – 1:50
and involvement in NYC's open data community. How did those two things come together for you?
Speaker 1
1:51 – 3:55
Yes. So it was it's an interesting story. So I was getting my master's. This was probably about 2019, 2020, etcetera, something like that. And I got an email through the education sphere about New York City Open Data Week. The New York City Open Data Platform in conjunction with beta New York City and the mayor's office of technology and innovation each year host their own conference all about New York City open data. And I was like, wow. This sounds awesome. So cool that they're just a bunch of different projects, talks, etcetera, using open data relating to New York City. And I was like, I think I wanna do this one day. So I had graduated, and I was still very interested in doing something for New York City open data week. So I gathered my friends who were Knicks fans, the basketball team. I said, hey. Let's do an analysis on the New York Knicks, and let's do it with the hope of presenting it at New York City Open Data Week. We do a lot of elbow grease. The project itself gets picked, and we present at New York City Open Data. And I then find out that they have this program called New York City Open Data ambassadors, where volunteers like myself would host maybe a few different introductory sessions per quarter for the general public on how to use New York City Open Data Platform. So I went through maybe a a four week course just meeting with different people from beta New York City in the mayor's office on the platform, how to teach it. And that's how I became an ambassador. And then I very quickly found out that my students need a little more oomph if I really wanted them to learn how to use r and statistics and analytics. And I found that relevance drives rigor. So that's how I decided to merge the world of New York City Open Data in my classroom. Oh, I like that phrase relevance
Speaker 0
3:55 – 4:18
drives rigor. Yes. That's quite good. Thank you. It's the truth as I've found. And we use that phrase open data. Came up a lot in in your answer to that last question. There's probably folks coming into this a bit fresh. Maybe that's a concept that hasn't been in their daily life and whatnot. When we use that term, what are we referring to, and why is open data so important?
Speaker 1
4:18 – 7:54
Wow. So that's an important question. Not only what is open data, but why is it important? Do I have time for a quick little history lesson? Well, go right ahead. I love I love the history lesson. Thank you. So part of the the talk that I give when I do the introductory classes for people so we go to New York back in, like, the late to mid eighteen hundreds, and there's this guy, Boss Tweed, who's mayor of New York City. And he's extremely corrupt. He runs the Democratic party at the time, a lot of nepotism, and he's really just hoarding a lot of the city's resources. New York gets fed up and they're like, hey. We don't want this anymore. So not only does boss Tweed get out of office, but they create what's now and was then called the city record. So anything that was happening in the city, like a new building was going up or new jobs, etcetera, were all being posted in the city record. So it's the first time that people were actually aware of what was going on in the city. We then get, like, a hundred years later into the seventies, nineteen sixties, nineteen seventies, and we get to the foil laws, freedom of information legislation, if I have my acronym right. So what the federal government and what a lot of state governments, including New York City, did was say, hey. You can now request government data. So it's the first time it's really I mean, the city record too, but it's the first time where data's open where you're like, hey. I've got a question. Let me get the data from the government. And while this is great, it the problem is is that you have to know what you're looking for. Because if you don't know what you're looking for, how are you gonna request it? So then we get to, I think, 2012 where we get the freedom of information law. I could get that exact a little wrong, but New York City says, hey. In perpetuity, we are going to make all data around New York, New York State and New York City open, which means that you no longer have to request it. Now we have it available, so you can pull it at any time, anywhere. I'd say that's the best way to describe open data. Data that you don't need any permission to use. You don't need a specific access or request. As long as you have access to the Internet, you can gather it. And if we go on the New York City Open Data portal today, there are 3,000 different datasets. I don't know how many billions of rows of data, and I think there's 5,000 unique columns. So trillions upon trillions of different data points that you can use at anytime, anywhere. But when you describe it like that, it sounds maybe like it's in a healthy place right now. Would you would you say that that's, like, kind of the state of things as we think about it today? Yeah. I think with New York City open data I mean, I I this may sound trivial, but the fact that you can access New York City open data from anywhere in the world really shows how powerful it is. You don't have to just be a New York City resident or be inside New York City. You can do it from South Korea or from any of the other 50 states. And, no, not everything is on there. A lot of the times that there's sensitive information that can't be, which is fair. But also the fact that let's say you're looking for a dataset and you can't find it. You can actually submit a request to New York City open data, say, hey. I would like to use this data. Is it available? And if they say yes, then you're the the trendsetter, and you're the reason why this dataset could be made available for other people to use.
Speaker 0
7:54 – 8:14
Because of that, I'd say it's in an incredibly healthy space right now. Part of your work has involved using a programming language called R to put packages together. For folks that are unfamiliar with the language, how would you describe R and the things that it's good at and useful for? So for me, I always describe it as love at first sight.
Speaker 1
8:15 – 9:35
I was first introduced a lot of people don't feel that way, including a lot of my students. But for me, it was love at first sight. So I first learned about art in 2019. I was taking a statistic course in my master's program. Had the option of taking neuroanatomy or statistics my first semester, and I thought that statistics would be more relevant early on. And then it was, from the ground running. R is an open source tool that is fantastic at statistical analysis and graphs as well. Graphing, not only statistics, but numbers, anything you really want. Now can it be as versatile as something like Python? Yes. You can do web scraping, etcetera, but it's really meant for statistical analysis, working with numbers, and working with data. Very influential in the academic community. And what makes it so powerful, like a lot of open source platforms, is that people can contribute. So if you have an idea let's say you're trying to do something in r and it's not possible yet. You could be the one that builds it and influence the whole r community. And the best part about it is that it's free. And so in my family, we have the saying, if it's for free, it's for me.
Speaker 0
9:36 – 9:40
That's, that that that's a good motto, especially, you know, in this economy.
Speaker 1
9:41 – 9:53
Yeah. And if you don't have like like, there are a lot of programs. Let's take Microsoft Excel, for example, where you have to pay for it. And so you can replace Microsoft Excel by replacing it with r.
Speaker 0
9:54 – 10:08
So you don't have to pay for it, and it has way more capabilities than Excel does. What was the moment that made you combine that enjoyment and, well, love at first sight that you found with R and the need for folks to access open data portals
Speaker 1
10:08 – 13:01
and decide, hey. You know what? Like, I need to do that open source thing and build a a package for this. Yeah. So it first of all, I never started my R journey. I was like, I am going to be a package developer. I actually was like, wow. This is something that are that's gonna be way too advanced for me no matter what level I am, and it wasn't something that I was aspiring for. What really happened was I saw a opportunity to fix a problem, and I jumped on it. I started teaching my at the master's program, the psychological resource master's program in the psych department in Brooklyn College, and they needed someone to teach students r. Here I go. I suggested that relevance drives rigor, and I was trying to teach students r with datasets that were not relevant whatsoever. I don't know if anyone or has listened that's listening has used R, but they come with preset datasets. For instance, the MT cars dataset, which is very popular, and it's all about cars, but none of my students like cars. So when we're talking about cylinders, number of cylinders, or how fast something can go or anything cars related, I mean, I'm putting them to sleep. So I say, how can I make my class more engaging because that's the only way that these students are really gonna learn? And my number one priority is for them to learn. So I said, what if I change the datasets? What if I start incorporating New York City Open Data into my teaching instead of using these arbitrary base r or other arbitrary datasets? Great idea. However, to connect to r at the time, there was a few different ways. One, we'd have to download the data from New York City open data portal, upload it to R, and that could be cumbersome, especially if you wanna use, like, daily data. For instance, the three one one dataset is updated daily. So if we wanna use the most up to date data, each week, we'd have to redownload and upload, etcetera. And now there's folders and naming, etcetera. So I wanted to get rid of that. The second way is we can connect with APIs, but my students were brand new at R, and I needed to make sure that I didn't lose them. So I didn't wanna take some time to just talk about APIs, talk about that computers, how computers talk to each other. I would be getting away from what I really wanted. So part of teaching R is teaching about packages and how to install them, their power. I said, screw it. Let's just make my own package. Let's see if I can make the New York City open data package. I can help myself and these nine students,
Speaker 0
13:01 – 13:32
and that's how it became. I feel like we have a joke in there somewhere about it being that, like, the car's dataset was is unpopular at a New York school. I feel like there's something in there about no one in New York has a car or something. Yeah. That that's a good I have a feeling we could work on that. So you mentioned, like, not wanting to teach kind of that, API interaction stuff. Like, how, toilsome was that for folks, like, before packages like that existed? Like, what kind of process is it to then try to kinda, like, do that yourself? Well,
Speaker 1
13:32 – 14:23
first, let me say that the New York City open data portal makes it extremely easy to access their API. Each different dataset comes with an API endpoint that you can copy and paste. So that is great by them. I'd say in r, there is the r Socrata package that you can use. You have to understand a little bit okay. Like, first, if I wanna use all these different datasets, I have to get all their different JSON links. Right? So endpoints. Excuse me. So I gotta make sure I keep track of that. And if I want to keep track of that and use it, then I have to use maybe the h t t r package or a few other packages. So it can be a little cumbersome. And if someone's trying to learn r in the beginning,
Speaker 0
14:23 – 14:46
it can get visually, it could look like, woah. This is a little out of my way. I see. Yeah. It sounds like maybe getting into, like, intermediate level kind of stuff, like dependency management. And then you kinda have to understand a little bit about HTTP request kind of stuff probably at that point, and you probably have to manage, like, API key kind of things or auth tokens or Yeah. Like, that's probably stuff that you don't wanna do when you're just getting past hello world with the language.
Speaker 1
14:47 – 14:56
Yeah. I really don't want my students to be even more overwhelmed, and that's something that I thought could push them over the edge. I recall from our prep conversation,
Speaker 0
14:57 – 15:23
that you got to work on a pack on the package at the city level. But as you were kinda working along, you discovered that there is a state portal with similar capabilities, which then led you to doing more building. I think you ended up with, like, three different packages if memory serves. So far. What was that moment of discovery? Like, can you maybe we can live a little bit vicariously through you. Yeah. So imagine that I've been working with the nearest city open data portal for a few years now. Very comfortable with it and totally
Speaker 1
15:24 – 17:30
I was in awe when I first started with it, and now I'm I'm so happy that it's been there, and I'm pretty advanced in it. So it was maybe the last day or the day before the last of the New York City open data week twenty twenty six, the big conference that they have every week. And there was a talk from New York City council data team members. I was like, wow. That seems so fun. I'd love to see how the city council is using data to influence policy. So I go, and all of a sudden, they pull up a different portal. It looks almost the same, but the coloring is a little different. And it's the New York State open data portal, and I am totally in shock. I couldn't believe it. I I had I was living under a rock thinking that there was only the New York City portal and not like a New York state or anything else. And as soon as I saw that, I was like, I have to do this too. It really being from New York and working in New York with CUNY, I was like, I have to make sure that as much data as possible can be made made available for my students or anyone else because New York City is part of New York State, and they talk, and they're somewhat one entity. So that's when I started to build the New York State open data package. Now in somewhat of a confusing manner, New York City has what's called the MTA, the Mass Transit Association. So buses, subways, etcetera. The MTA is not technically a New York City agency, and it works for New York State even though it just lives in New York City. And it kind of is on the New York State open data portal, but, like, on a kind of its own section. So I said, okay. What if I make the MTA open data package this way? It's kinda like the intermediate or bridge between New York City and New York State, which it kinda acts as literally. So now I have a a small little New York open data ecosystem.
Speaker 0
17:32 – 17:41
Were were there, were there any lessons you took from building of the first package that then kinda helped you out as you say built the second and the third?
Speaker 1
17:42 – 20:52
Yeah. So that's a fantastic question. When I first created the New York City open data package, I submitted it to our OpenSci, which is a community with the drive to make r as fantastic as possible. And it's really like a stamp of approval. So in the review process, when you submit a package junior to rOpenSci, there's a review of the package. And so you have an editor review it and then three independent reviewers, all with the intent of, hey. We're part of the art community. We wanna make this as strong as possible. And sometimes people are too close to the sun, and they don't even know that, hey. You can do something a little better, and that'd be either making your package more streamlined or easier to maintain or maybe impacting more people. So in my first iteration of the New York City open data package, I had almost 40 different functions. Here was my thought. There are when you go to the New York City open data portal, they have a list of the most viewed. So I said, okay. Let me just take the top whatever there were and make a function that pulls just that dataset. This way, if you were someone that didn't know what you were looking for and wanted to explore using r, you could see 40 different functions and be like, oh, like, I know what three one one is. I know what motor vehicle crashes is because it literally says that in the function name. From a reviewer standpoint and from a maintenance standpoint, terrible though. The fact that if you have to change something about your code, you have to change it in 40 different places means that you're gonna mess up. And sometimes, I'm not the most keen on attention to detail, so that could lead to one spelling error or one line of code being deleted, and now the whole package is in disarray. So, thankfully, now this iteration is we were able to take the metadata dataset on New York City Open Data that hosts all of the API endpoints metadata regarding it. There's over 3,000 different datasets if I believe, and all the information about that and use that to pull any data from New York City open data. So now there's three different functions instead of only 40. The first one is, hey. Let's get a list of any dataset that's available to you. The second is, hey. Using that list, put in the name that you want, and we can pull the dataset. And the third is just in case you found a dataset that is not on that meta list, you can put in the API endpoint, and it'll pull right in from you. So, thankfully, this was all done before I did the New York State and MTA open data packages because I would have wrote, I don't know, anywhere from a 100 to 200 maybe depending on the popularity of some. And that would have been from a maintenance standpoint. A nightmare and really not the best way to optimize the packages. Well, that that's such a good
Speaker 0
20:53 – 21:15
lesson. Actually, I think folks out here would probably do wanna learn from that story. It's very much like where do you put your layer of abstraction in the technical design of something kind of a thing? You know, it's like, you started with it at, like, my abstraction is about the datasets. Right? And now you shift it to it's about, like, the actions you take Yeah. Which is, like, a very interesting kind of refactor to to be doing.
Speaker 1
21:16 – 21:52
And a lot of the I didn't mention this, but it's it's pertinent to know. The package is really meant for myself and my students, but mostly my students. I didn't think that anybody else would want to use this, and it's it's crazy that other people have. But my thought was, okay. Like, my students have an interest in these datasets. Let's make it as easy as possible for them. So the first functions that I wrote in the package were just pulling their datasets that they had a desire for. And then it started with, alright. Let me use the most popular ones, and then it gets to 40. So is it less direct?
Speaker 0
21:52 – 22:20
Maybe. But it's definitely more streamlined, and there's way more opportunity now. And I think something that's important for folks to pull out of the story also is that you probably wouldn't have gotten to the place where you're thinking about, like, how could it be more streamlined if you didn't at least get to the point where you could pull some datasets. Yeah. So, like, if if you're just, like, starting a project and you're just I just I'm gonna make this function for this dataset just so you have something that works. That's, like, fine to start. You can always, like, make something better. Right?
Speaker 1
22:21 – 23:03
Yeah. I am a big believer in just start. Like, just get something down. I even remember when I was younger in school, and the hardest part for me to write an essay was just starting. But once I started, I was able to fly by. And to that exact point, my first dataset was just gonna be the 311. I was gonna see if I could just have a package with the 311 dataset. And I submitted it to CRAN, and they said, hey. Is this package done? There's only one function. Like, you could have one function, but we typically like if there's finished products. And I was like, alright. Like, let me do more then. And so it went from one to 40 to now
Speaker 0
23:03 – 23:53
3,000. I'd like to connect this a bit with, some of the education stories you've been talking about. You mentioned that MT cars thing where, you know, you discover, like, hey. I have this dataset. People aren't super into it. So then you you went on this adventure to kinda get more relevant datasets. That kinda, like, quest for relevant relevance, ended up putting your students in a place where there's a I believe you told me there's electronic book up Yeah. With, nine chapters each representing some sort of, like, interesting exploration of a question that they did. And as I kinda click through it, there's, like, a ton of breadth and variety, the types of questions that folks were interested in in, like, nerding out about or exploring. What are these projects like, and what does their diversity, say about the benefit of letting folks kinda follow their curiosity with the sort of learning?
Speaker 1
23:54 – 26:45
Fantastic question. So if I may jump back a little bit, I've been teaching for a good amount of time. And what I have found is that when you provide students an opportunity to be creative, they freeze. And it's unfortunate because they've got so many good ideas and there's so much opportunity and potential, and I really want them to explore it. Hey, You got a question? Go answer it. Let's have fun. It's okay. Let's make mistakes. If the question doesn't work out, no problem. We'll get another question. It's not that serious, and there's so much merit in exploration. And I think that's really what the book represents. So the final project was a research project, and the goal was to potentially present it at New York City open data week. I didn't know if they would be able to and or if it'd be good enough, if it fit the theme, but I wanted to for them to create something that could be presented. So I gave them two rules. One, it had to be about New York City, and two, it had to use open data. Any question that you had relating to that doesn't matter. Just has to fit those two criteria. And all nine of my students kind of freaked out a little bit. Professor Martinez, what should I do? What should my topic be? I don't know. What do you like to do? Oh, well, I like to I like to eat at restaurants. Okay. So let's answer a question about restaurants. What else do you like to do? Well, I like to go to museums. Okay. Can you create a relationship or a question that pairs the two? Well, let me see about that. And there is so much opportunity. And and when you look at my students' projects, it it maybe is the first time in their academic career where they get to look at something fun and interesting to them. And it allows them to hone the skill instead of just learning a tool. Because my whole class was learning r, and they were gonna have to learn r no matter what, if they had this research project or not. But now they got something that was tangible and excitable and relevant strives rigor. So now it's relevant to them. And so one of my students wanted to see if there was an impact on basketball players' performance and team performance, if they played at Madison Square Garden or not. Another one wanted to see if there was a relationship between mold complaints and domestic violence in New York City. Another one wanted to see if there were better restaurants surrounding museums in New York City. And so all beautiful and and eclectic topics. And it's it's amazing because they could be proud of it. They have told me that they went to their parents. Hey. Look what I did. This is so cool. I would describe it as refrigerator worthy instead of just another assignment. That's that's so cool. And
Speaker 0
26:45 – 27:11
something that that occurs to me with that is it's it's not really just as I think as you you were saying, it's not even really just about learning r at that point. Like, the by pushing, like, to get these questions about things you're interested in, They're kinda like learning how to basically design natural experiments. They're engaging with the scientific method, and those skills, whether using r or some other tool because, like, a job makes you one day, those skills are still relevant, in any of those contexts.
Speaker 1
27:12 – 27:41
Number one, think you can have fun with things, which I think is an underutilized motive. And number two, you're exactly right. It's let's build on the skills that we have learned, and let's build something creative and something that's relevant to you. Because if you're more interested in it, you're more likely to do a better product and have more fun with it and maybe not procrastinate as much, and you get to learn how
Speaker 0
27:41 – 28:17
you actually work best. As I looked through the the different chapters that folks had had made, I noticed that, the way they're put together, it's as though they're kinda designed to be readily reproducible. Like, I could take kind of this the code snippets they built there. If I pulled the same data, I could run it and go, oh, like, I found the same result. Like, I can see that, like, this thing you built is legit. Was that something that like, the way it's structured, that approach that was intentional, or did it kind of emerge as the the projects were coming along? Great question. No. That was the axiom of the entire project. So in a more condensed
Speaker 1
28:17 – 29:41
title, the class that I was teaching was reproducible research using r. The goal was to have a class that helps students learn what to do after they've done their experiment and they have their data. And I cannot stress how important reproducibility is in the scientific community because that's what validates your findings. You want your experiment to be reproducible. You want it where if someone runs it 100 times, a thousand times, 100,000 times, you get the same results. Because if I do it one time and I get something and I do it the exact same way I get another, that's not reproducible. That's not fact anymore. That's not something that we can utilize. You wouldn't want the the new vaccine or the new medicine to have different results every time they studied it. And so when my students are in which they're now completing their thesis, you want anyone in the scientific community to be able to replicate it so that they know it's legit. And so same thing with this. Sure. We're not working with a big company, and this is just for fun and a research project, but the underlying axiom is still true. We need to make sure that our work is reproducible. It's the most important part in my opinion. That does seem like a really
Speaker 0
29:41 – 30:12
critical lesson, particularly for students. Right? Because as you kinda go into life, there's many situations where the ability to then I guess if you make something that's reproducible, that maybe also gives you the skill of being able to know what to look for to see if you can reproduce something else. Right? Like, whether you're at work or you're reading about current events, there's often times you run across a paper or something and be great. Maybe you're curious and you, like, wanna try to test it yourself. And it so it seems like a valuable skill that folks are kinda learning as they build something that that allows for that.
Speaker 1
30:13 – 30:39
I I would like to think so and even reproducibly reproducibility within their own self. Like, okay. Let's write a comment to describe what this piece of code does. Because how many times, at least in my world, I've gone back to code that I thought was great, and I've looked at them and, like, what does this even do? And so even for even outside of other people using it, making sure it's reproducible within your own
Speaker 0
30:40 – 31:06
individual self. Oh, man. That what does that even do reaction? That's like me looking at my, like, six months ago code. And sometimes it's brutal. You're like, what did who wrote this? I didn't write this. When you, think about your experiences working with New York City and New York State's respective open data infrastructure, what are things that you're thinking are going well right now and things that you wish are would be different? Oh, great question.
Speaker 1
31:06 – 33:11
Number one, they I have to give them so much credit. First of all, very user friendly, and they really want people to access it. New York City open data portal has, I think maybe once a month, if not, once every few months, an introductory course. So that's what I teach as an ambassador, and it's for free. And you can come and learn how to use the basics. So imagine you didn't know what the New York City Open Data Platform was like. You could take this course, and you'll come out having a pretty good understanding of what it offers and what to do. I also have to give them credit in that very user friendly, and they whatever update they just made, the ability to use it has been updated and is so fast now. So before let's say I wanted to make a map. Right? I had three one one request, and I had the last, whatever, 100,000. I'd have to filter, so maybe there's only 500, or else my the rendering of my map would take so long. You'd think you'd get the circle of death just keeps spiraling and spiraling. Now almost instantaneous, which is a huge a huge impact on not only my teaching, but if people wanna create things on their own. I think if I had to make suggestion, one thing that has been that's available but I did not know is that they have this metadata package, the metadata dataset, which has the metadata on all the different datasets, what category they're in, how many views they have, downloads, etcetera. And that's not something that I was made readily available about had made known that was readily available until last month. So I'd have been an ambassador ambassador for too long, And so I think that could be promoted a little bit more because maybe that's something that other people can utilize and increase their potential to use the platform. Let's say there's some intrepid,
Speaker 0
33:12 – 33:36
civic minded person out there. Maybe they haven't had a chance to be in your class yet. And they're listening to this and they're like, man, I, you know, I wanna learn this stuff. I wanna I wanna learn how to use open data to answer questions. I wanna learn these kind of, like, research skills. What sort of advice would you give them as they just try to get started to kinda hearken back to our, you know, it's important to get started thing for them before?
Speaker 1
33:36 – 34:44
Say, number one, play around and make mistakes. I don't know how relevant this quote is, but yesterday, I was in the Posit data lab. They have, like, one, I think, once a month. And one of the pres presenters said, don't let AI take away your stupidity. Really meaning, like, just make some mistakes, have fun, play around. And so if you are an intrepid person that's trying to get involved, first, go on the New York City open data portal and try to answer a question. Even something so simple as who complains the most about potholes in each borough. So something like that. I think if you are a little more advanced and know how to use r or a different programming language, go use the New York City or New York State or MTA package and do the same thing. Explore, build a graph, something like that. Figure out a question that take a question you would like to see answered in your city and see if you can answer it in New York City. And have fun and reach out to me. I'm so interested in partnering with other people and exploring how much more this ecosystem can grow. I mean, that, pothole comment,
Speaker 0
34:45 – 35:41
takes me back a bit. Folks who've listened to this podcast for, like, a long time might remember that, there's we had an open, episode where someone was talking about, like, the Chicago open data portal. And it was very much one of those, like, you never quite know what impact the question you ask and use the data to do is gonna have where there's this whole thing where there's snowplow data. Mhmm. And folks ended up kinda finding a pattern where, wow, it's strange. Like, there's the side road that, like, the snowplow always seems to go to first even though there's, like, primary roads that aren't done yet. And, like, the algorithm is kinda like they're supposed to do these kind of thoroughfares, and there's, like, a process. Right? Mhmm. But for some reason, the street and then it turns out it was like an alderman's private residence street. It got, like, a bunch of trouble, and there's, like, a whole, like, news story thing about it. But, yeah. You know, you never really know what's gonna come of your curiosity, I guess, is the lesson from that kind of thing. Yeah. And I I bring up potholes that's that's fun. I'm interested in learning more about that because,
Speaker 1
35:43 – 36:11
Donnie and he's not the first mayor to do this, but they had a recent pothole blitz where they fixed, like, 7,000 different potholes across New York City. So I'd love to see if there is any relationship into kinda like what you were talking about, like, where the potholes specifically were fixed. And maybe in a perfect world, maybe they used New York City open data to find the potholes that were the most complained about, and maybe they fixed those. Maybe they didn't. Who knows?
Speaker 0
36:11 – 36:39
Who knows? Maybe one of your listeners will research it and find out. Oh, yeah. Yeah. And, actually, if one of those folks are out there, they should let me know if they if they build a little a little project for that because yeah. Because, I mean, you can even maybe go so far as to try to extrapolate, like, where should I think that there's the greatest pothole risk based on, like, past Yeah. Occurrences. You know? You get into odd things about microclimates and things or, like, where water flows through drainage, and I'm sure there's a lot of complex stuff that goes into how a pothole forms.
Speaker 1
36:39 – 37:02
Yeah. Even I'm thinking, like, you could probably match some traffic data on that as well because the more cars, the higher erosion most likely. For sure. Or or to your point, maybe you could do a relationship between the amount of cars and maybe flooding because they have flood complaints as well. See if there's a relationship between potholes, flooding, and or
Speaker 0
37:03 – 37:08
amount of traffic. Oh, there you go. I think we just spun up, like, three or four different project ideas for your future future students. Yeah.
Speaker 1
37:09 – 37:14
Who knows? Maybe in the next iteration, one of the next chapters is that question exactly. Oh, that'd be that'd be fantastic.
Speaker 0
37:16 – 37:29
And related to this kinda, like, getting involved kinda question thread, if there are folks out there that, maybe they're around New York or New York interested and they wanna get involved with open data New York City, how should they go about doing that? I would say
Speaker 1
37:30 – 39:00
reach out, go on GitHub, see what's available, and fork the repository. Play around, see what you can build. I would also say that, Ryan, you inspired me during our prep call last week because I say I had the same, like, oh, man moment when I got introduced to the New York State open data portal. When you said, man, it'd be so cool if there were other packages related to other city or state open data portals. And I was like, I didn't even think about the fact that other states and cities have the same thing. So So I'm currently trying to develop a Chicago open data package for our yeah. And I I have to again, a lot of the credit goes to the rOpen Science community for helping me out, but a lot of the code is transferable. So I've been kind of able to just plug in, like, the API JSON link, like, the main one and play around. And and right now, I'm in the testing phase. Hopefully, maybe even by today, later today, it'll be submitted to CRAN. But I'd love if people took some of the code that I have for any of the packages I have and tried to make it for their own portals. I have noticed a lot of the portals seem to be made by the same company. They look almost identical. Like, if you look at New York City open data platform and the Chicago one, kinda the same deal. Yeah. I suppose that makes sense. It's,
Speaker 0
39:01 – 39:10
not exactly it. I mean, the data is different, but it's, like, the same problem to solve it. Like, how do you, you know, take a bunch of different formats of data and reliably share it?
Speaker 1
39:11 – 39:21
I know that New York State and New York City have the same vendor. I it maybe that Chicago and maybe LA or Austin also have the same vendor. Who knows? Maybe there's one monopoly of
Speaker 0
39:21 – 39:47
people that are making open data platforms. Yeah. And it seems that we, we've ended up with, like, a surprise call to action, here here on the tail end where it's like, hey. Like, you know, reach out if you're interested in figuring out your own municipality or state's open data portal, and maybe if you maybe you wanna help put together some more package there. Also, it sounds like you're already kinda working on a Chicago one, so maybe the Chicago Chicagoans in the audience might wanna get involved with you to help out.
Speaker 1
39:48 – 40:01
Please. I'd love that. And if someone wants to help build it before I publish it, I'd gladly release myself of it, and someone can take it on. I'd love to partner with it with someone on it. Cool. And then I think for those things,
Speaker 0
40:02 – 40:38
maybe we can put some sort of contact info thing in the show notes for folks that wanna reach out and say like, hey. Hey, Christian. I wanna help you out with this stuff. Definitely. Cool. And on that note, Christian, thank you so much for for joining us here on CivicTickChat. This was a a really fun conversation, and I think folks will find, interesting stuff to bring into the day, interesting lessons they learned that they, might be able to bring into their work or hobbies or whatever you wanna call it. Yeah. I hope so. Thanks so much for having me. This was a lot of fun. Visit us on the web at civictech.chat, or subscribe to us for content updates wherever it is you download your podcasts.