https://developer.apple.com/swift/
This lowers the bar for entry to the iOS market.
Does it also lower the bar for Mac OS X?
Can it be used to write command-line command-line applications ("scripts")? It has a REPL, which means it can do a kind of "just-in-time" compile and run. This is how Python works, so perhaps this is a viable mode for using Swift.
Via the Objective-C and C compatibility, it has full access to the POSIX libraries, as well as Cocoa, so it can clearly be used to build command-line apps. It might lack the flexibility of Python, since it's compiled. But C (C++, Objective-C) with automated memory management is still a gigantic victory for writing fast and reliable programs.
Can it be plugged into Apache to write backend applications? It's compiled, and compatible with C and Objective-C. So, one can imagine that a mod-swift component in Apache might be possible. It might be better to work through existing FCGI interfaces and write stand-alone Swift back-ends. This would require a bunch of libraries for database API's, template rendering, request and response processing, and the various bits and pieces that make up a rich web development environment. But this is largely available for C and C++, making it available to Swift-based backends.
Is one language even a desirable goal?
The idea of having one official version of the class definitions seems very helpful for capturing knowledge and managing the intellectual property that is embodied in application logic.
Rants on the daily grind of building software. This has been moved to https://slott56.github.io. Fix your bookmarks.
Moved
Moved. See https://slott56.github.io. All new content goes to the new site. This is a legacy, and will likely be dropped five years after the last post in Jan 2023.
Thursday, June 19, 2014
Thursday, June 12, 2014
TDD, API Design and Refactoring
See this short discussion on a Stingray Reader feature:
https://sourceforge.net/p/stingrayreader/discussion/COBOL/thread/d2132851/?limit=25#2a3a
This turned into an exercise in pure TDD.
<rant>
I'm not a fan of applying TDD in a strict, death-march fashion.
I see the comments on Stack Overflow that indicate that some folks feel strongly that strict TDD is somehow helpful. While "test before code" is laudable and often helpful, there's no royal road to good software.
Design involves a great deal of back and forth between code and test. A great deal.
It's logically impossible to write a test without having thought about the code. In order to write the test first, there must be a notional API against which the test is written. Anyone who requires that the test file must be written before the notional class or module is just playing at petty tyranny.
The notional design -- the rough outline of the class or module -- can be written into a file before any tests. It's okay. It is still test-driven because the considerations of testability drove the design process.
In particular, when starting "from scratch" -- with nothing -- writing tests first is senseless. Some module or package structure must exist for the test modules to import.
</rant>
Having ranted, it still arises that the tests do come before any code under some circumstances.
In this case, the requested functionality was quite difficult to visualize. However, it was possible to cobble together a test case that simplified the problem down to something like this this:
In COBOL, the program would use logic like IF Header EQUALS "ABC" THEN MOVE Body TO ABC-Segment. We need a way to handle something like this in Python so that we can parse the EBCDIC COBOL data.
This summarized example allowed construction of a test case that made use of a API that might have existed. I was pretty sure I had a test case that showed an approach.
What Actually Happened
Since the application already had 178 unit tests, there was plenty of structure that worked.
The single new unit test relied on a notional API that wasn't really in place. The new test bombed grotesquely.
There are two solutions:
I started out chasing the second option. I tweaked some things. More tests failed. I tweaked some more things. The new test finally passed, but another test was failing.
Some careful study of the failing test revealed that my approach was wrong. Way wrong.
The notional API was a bad idea.
The tweaks to make it work were a worse idea.
Back to the Lab Bench
At this point, I had made enough changes that the only thing to do was copy the new test and use the Git Revoke on the local changes to unwind the awful mistakes.
Staring again, I had a slightly better grip on the relevant code. I had a failing test. I tried a different approach that wasn't quite so inventive. This meant modifying the test.
I actually went through a few iterations of the test, using the test method as a kind of lab bench.
A more Pythonic approach to the lab bench is to work from the >>> prompt. I think that all of the exemplary projects use the >>> prompt examples in their documentation. This is a way to narrow and clarify the API. As projects get big, they can sprawl. New features can wind up with many imports to pick and choose elements from existing modules.
When it becomes difficult to use the >>> prompt as the lab bench, that's a sign that the API is too complex. Refactoring must happen.
Using the unit test framework as the lab bench was a hint that something had drifted out of tolerance.
However. I did get a test which passed. Yay. Sort of.
The test code was hideous.
TDD and API Design
The point of TDD, however, is that we have a working suite of tests. Refactoring won't break anything.
The point was that the hideous API could be rewritten into something that both
The rule is this:
Sadly, Java doesn't have this kind of boundary. Java programming can spin into quite complex API's, limited only by the laziness of the programmer who avoids refactoring.
Or the malice of the programmer's manager in not allowing time to refactor.
https://sourceforge.net/p/stingrayreader/discussion/COBOL/thread/d2132851/?limit=25#2a3a
This turned into an exercise in pure TDD.
<rant>
I'm not a fan of applying TDD in a strict, death-march fashion.
I see the comments on Stack Overflow that indicate that some folks feel strongly that strict TDD is somehow helpful. While "test before code" is laudable and often helpful, there's no royal road to good software.
Design involves a great deal of back and forth between code and test. A great deal.
It's logically impossible to write a test without having thought about the code. In order to write the test first, there must be a notional API against which the test is written. Anyone who requires that the test file must be written before the notional class or module is just playing at petty tyranny.
The notional design -- the rough outline of the class or module -- can be written into a file before any tests. It's okay. It is still test-driven because the considerations of testability drove the design process.
In particular, when starting "from scratch" -- with nothing -- writing tests first is senseless. Some module or package structure must exist for the test modules to import.
</rant>
Having ranted, it still arises that the tests do come before any code under some circumstances.
In this case, the requested functionality was quite difficult to visualize. However, it was possible to cobble together a test case that simplified the problem down to something like this this:
01 Some-Record.
05 Header PIC XXX.
05 Body PIC X(17).
01 ABC-Segment.
05 Field-ABC PIC X(17).
01 DEF-Segment.
05 Field-DEF PIC X(17).
In COBOL, the program would use logic like IF Header EQUALS "ABC" THEN MOVE Body TO ABC-Segment. We need a way to handle something like this in Python so that we can parse the EBCDIC COBOL data.
This summarized example allowed construction of a test case that made use of a API that might have existed. I was pretty sure I had a test case that showed an approach.
What Actually Happened
Since the application already had 178 unit tests, there was plenty of structure that worked.
The single new unit test relied on a notional API that wasn't really in place. The new test bombed grotesquely.
There are two solutions:
- Modify the test.
- Fix the notional API so that it works properly.
I started out chasing the second option. I tweaked some things. More tests failed. I tweaked some more things. The new test finally passed, but another test was failing.
Some careful study of the failing test revealed that my approach was wrong. Way wrong.
The notional API was a bad idea.
The tweaks to make it work were a worse idea.
Back to the Lab Bench
At this point, I had made enough changes that the only thing to do was copy the new test and use the Git Revoke on the local changes to unwind the awful mistakes.
Staring again, I had a slightly better grip on the relevant code. I had a failing test. I tried a different approach that wasn't quite so inventive. This meant modifying the test.
I actually went through a few iterations of the test, using the test method as a kind of lab bench.
A more Pythonic approach to the lab bench is to work from the >>> prompt. I think that all of the exemplary projects use the >>> prompt examples in their documentation. This is a way to narrow and clarify the API. As projects get big, they can sprawl. New features can wind up with many imports to pick and choose elements from existing modules.
When it becomes difficult to use the >>> prompt as the lab bench, that's a sign that the API is too complex. Refactoring must happen.
Using the unit test framework as the lab bench was a hint that something had drifted out of tolerance.
However. I did get a test which passed. Yay. Sort of.
The test code was hideous.
TDD and API Design
The point of TDD, however, is that we have a working suite of tests. Refactoring won't break anything.
The point was that the hideous API could be rewritten into something that both
- Passed all the tests, and
- Was usable at the >>> prompt.
The rule is this:
If the API doesn't make sense at the >>> prompt, it's incomprehensible
Or the malice of the programmer's manager in not allowing time to refactor.
Thursday, June 5, 2014
Grace Murray Hopper
Read this: https://www.indiegogo.com/projects/born-with-curiosity-the-grace-hopper-documentary#home
Consider donating to preserve history.
Read this: http://www.i-programmer.info/news/82-heritage/7368-crowd-fund-film-about-grace-hopper.html
Her legacy is often overshadowed by folks like Bill Gates and Steve Jobs.
Consider donating to preserve history.
Read this: http://www.i-programmer.info/news/82-heritage/7368-crowd-fund-film-about-grace-hopper.html
Her legacy is often overshadowed by folks like Bill Gates and Steve Jobs.
Thursday, May 29, 2014
Stingray 4.4 Update -- the Posix split command applied to COBOL files
Here's an interesting problem. Implement the split command for mainframe COBOL EBCDIC files with their BDW and RDW headers.
The conventional split can't handle COBOL EBCDIC files because they don't have sensible \n line breaks. Translating an EBCDIC file to ASCII is high-risk because COMP and COMP-3 fields will be trashed by the translation.
The itertools.groupby() function can break a record iterator into groups based on some group-by criteria. We can use this to break into sequential batches.
This expression will break the iterable, reader, into groups each of which has a size of batch_size records. The last group will have total%batch_size records.
The with statement allows us to make each individual group into a separate context. This assures that each file is properly opened and closed no matter what kinds of exceptions are raised.
Here's a typical script.
There are several possible variations on the construction of the reader object.
The conventional split can't handle COBOL EBCDIC files because they don't have sensible \n line breaks. Translating an EBCDIC file to ASCII is high-risk because COMP and COMP-3 fields will be trashed by the translation.
If the files include Occurs Depending On, then the FTP transfer should include the RDW/BDW headers. The SITE RDW (or LOCSITE RDW) are essential. It's much faster to include this overhead. Stingray can process files without the headers, but it's slower.There are two essential Python techniques for building file splitters than involve parsing.
- The itertools.groupby() function.
- The with statement.
Along with this, we need an iterator over the underlying records. For example, the stingray.cobol.RECFM subclasses will parse the various mainframe RECFM options and iterate over records or records+RDW headers or blocks (BDW headers plus records with RDW headers.
The itertools.groupby() function can break a record iterator into groups based on some group-by criteria. We can use this to break into sequential batches.
itertools.groupby( enumerate(reader), lambda x: x[0]//batch_size )
This expression will break the iterable, reader, into groups each of which has a size of batch_size records. The last group will have total%batch_size records.
The with statement allows us to make each individual group into a separate context. This assures that each file is properly opened and closed no matter what kinds of exceptions are raised.
Here's a typical script.
import itertools
import stringray.cobol
import collections
import pprint
batch_size= 1000
counts= collections.defaultdict(int)
with open( "some_file.schema", "rb" ) as source:
reader= stringray.cobol.RECFM_VB( source ).bdw_iter()
batches= itertools.groupby(enumerate(reader), lambda x: x[0]//batch_size):
for group, group_iter in batches:
with open( "some_file_{0}.schema".format(group), "wb" ) as target:
for id, row in group_iter:
target.write( row )
counts['rows'] += 1
counts[str(group)] += 1
pprint.pprint( dict(counts) )
There are several possible variations on the construction of the reader object.
- cobol.RECFM_F( source ).record_iter() -- result is RECFM_F.
- cobol.RECFM_F( source ).rdw_iter() -- result is RECFM_V; RDW's have been added.
- cobol.RECFM_V( source ).rdw_iter() -- result is RECFM_V; RDW's have been preserved.
- cobol.RECFM_VB( source ).rdw_iter() -- result is RECFM_V; RDW's have been preserved; BDW's have been discarded.
- cobol.RECFM_VB( source ).bdw_iter() -- result is RECFM_VB; BDW's and RDW's have been preserved. The batch size is the number of blocks, not the number of records.
This should allow slicing up a massive mainframe file into pieces for parallel processing.
Thursday, May 22, 2014
Python Package Design, Refactoring and the Stingray Reader Project
We'll be digging into Mastering Object-Oriented Python. Chapter 17, specifically.
We'll also be looking at a big refactoring of the Stingray Schema-Based File Reader.
We can identify three species of packages.
One common design is a Simple Package. A directory with an empty __init__.py file. This package name becomes a qualifier for the internal module names. The package is simply a namespace for modules. We’ll use the package with something like this:
Another common design is the Module-Package. This is a package which appears to be a module. It will have a larger and more sophisticated __init__.py that is a effectively, a module definition. There are two variations on this theme. Sometimes we'll use this during the early stages of development because we don't know if the package will get really big or stay small. If we start out with a package and all the code is in the __init__.py, we can refactor down to a module.
The more common use for a module-package is to have the __init__.py import objects or other modules from the package directory. Or, it can stand as a part of a larger design that includes the top-level module and the qualified sub-modules. We’ll use the package with something like this:
or perhaps
The third common pattern is a package where the __init__.py selects among alternative implementations. The os module is a good example of this. We’ll use the package with something like this:
Knowing that it did something roughly like the following for us.
Refactoring Module to Package
The Stingray angle on this is the need to add iWork '13 numbers to the collection of spreadsheets which it can parse. The iWork '13 format is unique.
Previously, all of the spreadsheets fell into three families:
We'll also be looking at a big refactoring of the Stingray Schema-Based File Reader.
We can identify three species of packages.
One common design is a Simple Package. A directory with an empty __init__.py file. This package name becomes a qualifier for the internal module names. The package is simply a namespace for modules. We’ll use the package with something like this:
import package.module
Another common design is the Module-Package. This is a package which appears to be a module. It will have a larger and more sophisticated __init__.py that is a effectively, a module definition. There are two variations on this theme. Sometimes we'll use this during the early stages of development because we don't know if the package will get really big or stay small. If we start out with a package and all the code is in the __init__.py, we can refactor down to a module.
The more common use for a module-package is to have the __init__.py import objects or other modules from the package directory. Or, it can stand as a part of a larger design that includes the top-level module and the qualified sub-modules. We’ll use the package with something like this:
import package
or perhaps
from package import thing
The third common pattern is a package where the __init__.py selects among alternative implementations. The os module is a good example of this. We’ll use the package with something like this:
import package
Knowing that it did something roughly like the following for us.
import package.implementation as package
Refactoring Module to Package
The Stingray angle on this is the need to add iWork '13 numbers to the collection of spreadsheets which it can parse. The iWork '13 format is unique.
Previously, all of the spreadsheets fell into three families:
- CSV or CSV-like. Simple text.
- XLS. We relied on https://pypi.python.org/pypi/xlrd/0.9.3 to make sense of .XLS files.
- XML-based. We parsed the XML and located sheets, rows and cells.
iWork '13 uses Snappy compress and Protobuf Serialization. Without some documentation, the files would be incomprehensible. Read this: See https://github.com/obriensp/iWorkFileFormat. Brilliant.
The previous releases of Stingray had a single, large module to handle a variety of workbook formats. Folding iWork '13 into this module would have been lunacy. It was already large to the point of being painful to understand.
The original module will be transparently turned into Module-Package. The API (import stingray.workbook or from stingray.workbook import SomeClass) will remain the same.
However.
The implementation will involve a package with each workbook format as a separate module inside that package. At the top, the __init__.py will include code like the following.
from stingray.workbook.csv import CSV_Workbook
from stingray.workbook.xls import XLS_Workbook
from stingray.workbook.xlsx import XLSX_Workbook
from stingray.workbook.ods import ODS_Workbook
from stingray.workbook.numbers_09 import Numbers09_Workbook
from stingray.workbook.numbers_13 import Numbers13_Workbook
from stingray.workbook.fixed import Fixed_Workbook
This has the advantage of allowing us to include additional parsing cruft in each module that's not part of the exposed API in the workbook package.
The Mastering Object-Oriented Python book has more details on this kind of design.
The Mastering Object-Oriented Python book has more details on this kind of design.
Thursday, May 15, 2014
Want a copy of Mastering Object-Oriented Python? Free?
Want a copy free? See this contest: http://www.blog.pythonlibrary.org/2014/05/12/ebook-contest-win-a-free-copy-of-mastering-object-oriented-python/.
If you're really interested, I can sign a copy. That will double the shipping cost, so perhaps that's not the best idea.
The bad news is that the errata have started to trickle in. Some of the default serialization cases in chapter 9 aren't handled properly. The demonstrations don't exercise the defaults very well, so things happen to work for me, but don't work when generalized or pulled out of context.
Sigh.
What's important -- I guess -- is that we have a critical mass of readers who are applying the concepts and finding problems.
If you're really interested, I can sign a copy. That will double the shipping cost, so perhaps that's not the best idea.
The bad news is that the errata have started to trickle in. Some of the default serialization cases in chapter 9 aren't handled properly. The demonstrations don't exercise the defaults very well, so things happen to work for me, but don't work when generalized or pulled out of context.
Sigh.
What's important -- I guess -- is that we have a critical mass of readers who are applying the concepts and finding problems.
Thursday, April 24, 2014
Literate Programming with pyWeb 2.3
Updates completed. See https://sourceforge.net/projects/pywebtool/ and http://pywebtool.sourceforge.net.
The list of changes is extensive.
However, the essential API and the markup language for creating literate programs hasn't (significantly) changed. A few experimental features were replaced with a first-class implementation.
The interesting (to me) bit is this sequence of events.
I started out using Leo and Interscript as a literate programming tools. They worked. But they were larger and clunky and I wasn't happy.
I wrote my own too, not really getting the use cases.
I found pyLit and liked it a lot. For a long time, I liked it better than my own pyWeb tool.
Then I ran across some problem domains for which pyLit didn't work out well. It's not that I've abandoned pyLit, but I believe I'll focus more on pyWeb.
The Awkward Problem Domains
Here are the two awkward problem domains.
The list of changes is extensive.
However, the essential API and the markup language for creating literate programs hasn't (significantly) changed. A few experimental features were replaced with a first-class implementation.
The interesting (to me) bit is this sequence of events.
I started out using Leo and Interscript as a literate programming tools. They worked. But they were larger and clunky and I wasn't happy.
I wrote my own too, not really getting the use cases.
I found pyLit and liked it a lot. For a long time, I liked it better than my own pyWeb tool.
Then I ran across some problem domains for which pyLit didn't work out well. It's not that I've abandoned pyLit, but I believe I'll focus more on pyWeb.
The Awkward Problem Domains
Here are the two awkward problem domains.
- Historical Story Lines. In some cases, we want to describe a module or package based on the path of exploration. Rather than simply drop the design, we want to show the path followed which lead to the design. This can be helpful for certain kinds of pedagogical exercises where we're steering the reader through a process.
- Complex Packages that Don't Follow Python's Presentation Order. In some cases, we need to present things out of order. Python constrains us to have docstring and imports first. Our class definitions must proceed in "dependency" order. But this may not be the best order for explanation. Sometimes, we want to start with the "def main():" function first to explain why a class looks the way it does.
PyWeb handles these nicely. One of the handiest things is this for out-of-order presentation.
@d Some Class... @{
class TheClass:
this class uses the following imports
@}
@d Imports...@{
import this
import that
@}
We can then scatter imports through the documentation in the relevant places. And they follow the more interesting material.
When it comes to final assembly, we have this.
@o some_module.py @{
@<Imports for this module@>
@<Some Class that does the real work of this module@>
@}
This builds the module, tangling the imports into one cluster up front, and putting the class definition later.
Subscribe to:
Posts (Atom)